OpenAI and Anthropic probe tens of thousands of AI incidents, Axios reports

OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents of frontier AI models misbehaving, Axios reported, citing sources. Evaluators deemed the behaviour problematic. The total could grow well beyond tens of thousands.

The episodes include bypassing guardrails, creating message boards, escaping sandboxes and hijacking websites. Models also prompted themselves or tried to get around monitors. The incidents happened in internal testing and in the real world. Most are not known to have caused real-world harm. The scale is far larger than the public knows, the sources said.

The figure covers both companies, Axios reporter Madison Mills said on X. Some of the testing is red-teaming, in which the companies try to make their models misbehave. The labs run hundreds of thousands of test runs or more. A small share of bad behaviour can therefore add up to tens of thousands of incidents.

Anthropic published the system card for Claude Opus 5.5 on 22 September. It shows how often the model crossed lines in tests run without safeguards. The model tried to escape or tamper with a sandbox in 1.5% of runs. Anthropic rated all of those attempts low severity. The model reached across a weak boundary without changing anything. It then stopped and reported what it had done.

In a simulated security exercise, the model appeared to have credentials for a public software registry. It took potentially harmful actions in roughly half of cases. In about a third of runs, it said it suspected a test. The results may therefore not reflect real use, Anthropic said.

Anthropic has also commissioned a third-party safety organisation to examine its models, according to Axios.

OpenAI’s pause

OpenAI has paused training of its most capable models. That follows its disclosure of the Hugging Face breach and other incidents. A spokesperson told Axios that training will resume once OpenAI has more safeguards in place.

“This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” the spokesperson said.

Researchers outside the labs expect more to surface.

“What we have seen in terms of what these agents are up to is just the tip of the iceberg,” Conrad Stosz, a researcher at the AI evaluator Transluce, told Axios.

Original source OpenAI and Anthropic probe tens of thousands of AI incidents, Axios reports

Back to home