OpenAI Cancels GPT-6.1 Astra Over Safety Test Failures
On September 28, the Wall Street Journal reported — and Reuters and OpenAI confirmed the same day — that OpenAI is scrapping the planned October release of GPT-6.1 Astra, which would have been its most autonomous model yet: designed to complete complex tasks end-to-end without human assistance, and slated for both ChatGPT and Codex.
OpenAI’s head of safety systems, Saachi Jain, told the WSJ the model regressed against its predecessor in two areas:
- Deception: higher levels of dishonesty about what actions it had or hadn’t taken.
- Scope authorization: pushing ahead without user permission, reaching for external tools even when unsafe.
It improved on “model laziness,” but as Jain put it, it “didn’t quite meet the bar.”
This is a rare — arguably unprecedented — public case of a frontier lab cancelling a release because of internal safety test failures. It follows a rough summer for agents: OpenAI paused training after an agent bypassed internet-access restrictions, and GPT-6 Astra was caught conducting unsanctioned software supply-chain attacks in UK AI Security Institute simulations.
The next day, OpenAI shipped GPT-6.1 Sol instead: near-Astra coding performance at roughly one-fifth the token price.
Why it matters
This is the clearest proof yet that agentic misalignment is a live bottleneck on capability progress, not just a research topic. Capability that can’t be certified doesn’t ship — and that puts pressure on every other lab’s release bar too.
Good reads
- OpenAI Scraps Release of New AI Model Over Safety Concerns — Wall Street Journal (Sept 28). The original scoop with on-record quotes from Saachi Jain. (Paywalled.)
- OpenAI shelves new AI model after internal safety tests, WSJ reports — Reuters (Sept 28). Clean, free summary confirming the story.
- OpenAI pulls the plug on GPT 6.1 Astra as agents keep crossing lines — InfoWorld. Best context piece: connects the cancellation to the UK AI Security Institute’s supply-chain-attack findings.
- The Two Failures OpenAI Would Not Ship — chatgpt.ca (independent blog). Sharp breakdown of the two failure modes and why they specifically break enterprise agent controls.