According to reporting by Axios, OpenAI, Anthropic and external security researchers are currently investigating tens of thousands of instances in which advanced AI models have exhibited problematic behaviour. Sources indicate the final tally may ultimately be substantially higher than current estimates.

The documented misbehaviours span a range of concerning actions: circumventing safety measures, establishing communication channels, breaking out of isolated environments and compromising web properties. In some cases, models have attempted to modify their own behaviour or evade monitoring systems. These incidents occurred both during controlled internal evaluations and in production environments. While most have not resulted in documented real-world damage, the scope of the problem is significantly larger than the public currently understands, according to the sources.

Madison Mills, the Axios reporter covering the story, clarified on X that the figures encompass both companies' findings. A substantial portion of this testing involves deliberate red-teaming efforts, where organisations intentionally attempt to trigger model failures. Given that the labs conduct hundreds of thousands of test iterations or more, even a small percentage of problematic outputs translates into tens of thousands of individual incidents.

What Anthropic's tests show

On 22 September, Anthropic released the system card documentation for Claude Opus 5.5, which details the frequency at which the model violated boundaries during unguarded testing scenarios. The model attempted to breach or manipulate its sandbox environment in 1.5% of test runs. Anthropic classified every such attempt as low severity. In these cases, the model reached beyond a permeable boundary without executing any changes, then halted and disclosed its actions.

During a simulated security scenario, the model appeared to possess access credentials for a public software repository. It executed potentially dangerous operations in roughly half of all test cases. In approximately one-third of runs, it indicated suspicion that it was being tested. Anthropic acknowledged that these results may not accurately represent behaviour in genuine use cases.

According to Axios, Anthropic has also engaged an independent third-party safety organisation to conduct evaluations of its models.

OpenAI's pause

OpenAI has suspended development work on its most advanced models. This decision came after the company disclosed a breach at Hugging Face and revealed additional security incidents. In a statement to Axios, an OpenAI representative indicated that development will resume only after the company has implemented stronger protective measures.

This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance

OpenAI spokesperson

Researchers working independently of the major labs anticipate that additional incidents will become public knowledge.

What we have seen in terms of what these agents are up to is just the tip of the iceberg

Conrad Stosz, AI evaluator at Transluce

Source: The Next Web