OpenAI has cancelled the planned October release of its GPT-6.1 Astra model after testers raised safety concerns, the company told the , in the latest fallout from heightened scrutiny of the security implications of powerful model releases in recent weeks.
OpenAI had planned to release the model sometime in October, but testers found it performed poorly on alignment, a way of measuring its tendency to carry out unexpected or unintended actions, OpenAI head of safety systems Saachi Jain told the paper.
The model showed higher levels of deception than its predecessor, GPT-6 Astra, a measurement of a model’s transparency about disclosing which actions it took or didn’t take.
Unexpected actions
The model also showed issues with “scope authorisation”, indicating it tended to push ahead with tasks without asking for user permission, and would at times try to access external tools or services even if doing so might overstep security boundaries.
The decision came as OpenAI has begun reviewing training activity involving its models, resulting in the discovery and disclosure of dozens of new security incidents, including an unintended hack of Australia’s government-operated healthcare system that resulted in a breach of non-public data.
AI companies began paying closer attention to such incidents after AI platform Hugging Face in July disclosed that it had been hacked by autonomous AI agents, which later proved to be agents being tested by OpenAI.
Such concerns are in addition to the problem of users intentionally making use of powerful AI systems to carry out security breaches.