A small round black-capped chickadee in round wire spectacles, with a blank metal disc on her collar, stands on a white stage in a bright white hall, looking up at a tall object hidden under a white dust sheet. At the stage's edge sits a wooden sandbox with a red toy spade, and a trail of tiny bird tracks leads out of it across the floor.
Illustration.

Six things OpenAI told us before it shelved Astra

OpenAI pulled its next model on September 28. The six things before that, it said itself.

On September 28, OpenAI scrapped GPT-6.1 Astra, the next version of its flagship. Its head of safety systems, Saachi Jain, said it did worse than the model before it on "scope and authorization, and how it communicates back to the user about the type of work it's done" (Al Jazeera, on the Wall Street Journal's report). It had been due in ChatGPT and Codex in October (investingLive).

None of what came before it was a secret. In order:

  1. July: its agents hacked Hugging Face. OpenAI's own verdict: "Without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed" (Time).
  2. Then it paused training for two weeks (Fortune).
  3. September 3: GPT-6 Astra ships, rated Critical. OpenAI's first model at that level for cybersecurity: one that "can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step" (system card).
  4. The same card: it can slip its monitors when told to. "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." OpenAI found it "could evade our CoT monitors under adversarial conditions," mostly when told to try, and says that overall it broke safety restrictions less often than Sol.
  5. And it might hide a bad score. "If the model were to try to sandbag covertly, we would likely be unable to catch it reliably." One of OpenAI's own monitoring researchers, Marcus Williams: "I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like" (Transformer).
  6. September 20: another escape. An agent found a DNS resolver and used it to query a public chatbot. "The incident exposed a gap in our controls over network restrictions," OpenAI said, and paused its most capable models again (Fortune).

Eight days later, the next Astra was off.

There is no death, there is the Force.