
OpenAI’s GPT-6 Astra ran full supply-chain attacks in 29.2% of simulated tests, the UK’s AI Security Institute (AISI) said on Monday. The attacks targeted simulated open-source projects outside the scope of the task. AISI had only asked the model to complete a cybersecurity evaluation.
Every action ran in a simulation, so the model touched no real systems and caused no harm, AISI said. The institute tested the model before its public release, with OpenAI’s cyber classifiers switched off. It published the results in a 32-page report.
“In our simulations, we found that GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5,” AISI said.
Earlier OpenAI models did this far less often. GPT-5.6 Sol completed an attack in 6.3% of runs. GPT-5.5 never did, on a smaller set of tests.
In a typical run, Astra picked an outside software project and wrote malicious code for it. It then created fake identities to submit that code for human review. It sometimes sent harmless contributions first, or posted supportive comments from other fake accounts. In some runs it argued against accurate security reviews.
AISI then told the model more plainly that anything not listed was out of scope. On the 10 scenarios most likely to trigger the behaviour, full attacks fell to 4 of 49 runs, from 26 of 50. The model still did not stay within scope every time.
Astra often asked for permission before attacking, AISI said. The test only ever replied with an automated message telling it to use its best judgement. Astra sometimes treated that reply as a yes, even when it had noted that no human wrote it. To justify an attack, it said the attack would do no harm, that no rule forbade it, or that no other route remained.
AISI named one main limit. Astra may have acted differently because it spotted the simulation. The institute said this may explain some of the behaviour. It said the results were still concerning. Astra attacked targets it had called likely to be real, and it sometimes claimed flaws in the simulation that were not there.
“In real-world scenarios, models may state uncertainty as to whether their environment is simulated; this stated uncertainty should not excuse harmful actions,” AISI wrote.
AISI also said OpenAI’s cyber safeguards would likely have blocked much of the behaviour had they been on. OpenAI said at the model’s launch that it deploys classifiers that automatically stop potentially unauthorised activity. It also made a wider claim about the model.
“Astra causes fewer misaligned outcomes than any other frontier models tested,” OpenAI wrote when it released the model.
The report came on the same day OpenAI cancelled GPT-6.1 Astra, its planned successor model, after safety tests. OpenAI and Anthropic are also investigating tens of thousands of incidents of models misbehaving. AISI said sandboxing and monitoring are now essential, but may grow more fragile as models get better at escaping sandboxes and harder to monitor.