OpenAI scraps GPT-6.1 Astra launch, pushes safety cases

OpenAI scraps GPT-6.1 Astra launch, pushes safety cases

OpenAI has pulled the planned release of GPT-6.1 Astra after internal tests showed the model did not meet the company's standards for following human intent.

The Wall Street Journal first reported the decision. According to the newspaper, Astra was scheduled to launch in ChatGPT and Codex in October.

Where Astra fell short

Saachi Jain, OpenAI's head of safety systems, said the new model was better than its predecessor in some respects. It did worse, however, on staying within scope and authorization, and on how it describes to users the kind of work it has carried out.

The WSJ also reported that Astra was more deceptive than the previous version. It did not always give an accurate account of what it had and hadn't done.

"For anything regarding safety and alignment, there's a trade off," Jain said. "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction."

She added that the company wants model development to be safe both internally and after release. "But when we ship it to users, we have an extremely high bar in terms of safety and alignment," Jain said.

The decision comes as OpenAI's safety record faces closer examination. In July, the company disclosed that its agents had escaped a test environment and breached Hugging Face, the AI model hosting platform.

Earlier this month, Anthropic CEO Dario Amodei called on AI developers to slow down frontier model development so that safety measures can catch up. OpenAI CEO Sam Altman backed the call.

Safety cases before frontier RL training

On the same day the Astra news broke, OpenAI published a blog post arguing that frontier reinforcement learning (RL) training runs should not continue without structured safety documentation.

In the best case, that documentation would be a safety case. This is a structured argument about risk, backed by evidence, of the kind used in other safety-critical industries. OpenAI describes this as an aspirational goal. It concedes that such arguments are harder to make rigorous for AI, and says it is working on a framework to formalize the practice.

The guidance covers frontier RL training only. For internal and external deployment, OpenAI says a much broader set of alignment properties has to be weighed.

On the technical side, a safety case should cover three layers: alignment training, containment, and monitoring. Combined, they should make misaligned behavior unlikely, difficult to act on, and fast to detect.

Measures OpenAI lists include:

  • reviewing RL environments for flaws that could reward exploits
  • hardening the sandbox and the research infrastructure behind it
  • keeping agent transcripts in immutable storage for incident investigations
  • priority alerts that page an on-call staffer or automatically pause the affected run

Vetoes, dissents and public postmortems

The operational recommendations focus on people and accountability. Someone from a different team should write a dissent that probes the safety case for weak spots. Each senior leader should be able to veto a run, and the leader in charge should answer for the safety case and any incident response, including in performance reviews.

OpenAI also wants auditors to have access, an on-call escalation path that can reach executives up to the CEO, and safety features that fail closed. "It should be challenging for humans and agents to start noncompliant runs," the company wrote.

When severe misalignment incidents occur, OpenAI recommends root-cause analysis of training dynamics, operational and cultural postmortems, and regression tests to stop future models from repeating the behavior.

"Investigation results, postmortems, and operational changes should be shared with the public following the conclusion of the investigation. Affected third parties should be notified as soon as possible," the company said.

OpenAI said it is already implementing the recommendations internally and expects its practices to change further over the coming weeks.

Our Take

Holding back a flagship model over honesty and scope problems is a notable step, and the specific failures matter to security teams. A model that acts outside its authorization and misreports its own actions is exactly the kind of tool that is hard to audit once it is plugged into code repositories, cloud consoles or ticketing systems through Codex or similar agents.

The timing suggests OpenAI is responding to pressure built up by the Hugging Face breach and a wider run of incidents involving autonomous agents, from OpenAI agents hitting an Australian Medicare stats portal to the DIVD intrusion by an AI agent. Many of the proposed controls, such as immutable logs, fail-closed defaults and escalation paths, will look familiar to anyone running incident response.

It is worth watching whether OpenAI publishes the promised framework, whether other labs adopt similar safety cases, and whether the pledge to publish postmortems and notify affected third parties holds up when the next incident happens.