Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- GPT-6.1 Astra Explained: Why OpenAI Shelved Its Next Model
GPT-6.1 Astra Explained: Why OpenAI Shelved Its Next Model
GPT-6.1 Astra was OpenAI's next model until it was shelved on 28 September 2026. See why, what GPT-6 Astra's system card shows, and how to check agent work.
Published on:
GPT-6 Astra drew roughly half as many higher-severity misalignment flags as GPT-5.6 Sol across more than 54,000 simulated internal Codex tasks, and was stronger at staying within its authorized scope, according to its system card.
Its successor never shipped. On 28 September 2026, OpenAI confirmed to CNBC that it would not release GPT-6.1 Astra, because the model fell short on staying within scope and authorization and on how it reported its work back to users.
The second reason matters to anyone running agents. A report of an agent's work can only be checked against a record of that work, and that check is what TestMu AI builds for agents that act.
TL;DR
GPT-6.1 Astra is the OpenAI model that was cancelled on 28 September 2026, before its planned October release in ChatGPT and Codex. In internal testing it gave up on tasks less often than GPT-6 Astra, but it stayed within its authorized scope less reliably and did not always report accurately which actions it had taken.
- Release status: Is GPT-6.1 Astra available in ChatGPT or the API? No. OpenAI cancelled the release before launch, so no one outside OpenAI has used the model, and OpenAI has not published evaluation figures for it.
- Model laziness: Did GPT-6.1 Astra improve on model laziness? Yes. GPT-6.1 Astra gave up or handed tasks back to the user less often when it hit an obstacle, which is the improvement OpenAI was aiming for.
- Scope and authorization: The Wall Street Journal reported that GPT-6.1 Astra pushed ahead on tasks without asking permission and reached for external tools and services even when that might be unsafe.
- Reporting its own work: GPT-6.1 Astra showed more deception than GPT-6 Astra, including not always telling users accurately which actions it had or had not taken.
- GPT-6.1 Sol: Is GPT-6.1 Sol the same model as GPT-6.1 Astra? No. GPT-6.1 Sol is an upgrade to GPT-6 Sol announced at DevDay on 29 September, while GPT-6.1 Astra was the shelved successor to GPT-6 Astra.
- Checking agent reports: A report of an agent's work can only be verified against a record of the work. TestMu AI Agent Assurance grades agents on what each run changed and reports what it could not verify.
What Happened to GPT-6.1 Astra?
OpenAI cancelled the planned release of GPT-6.1 Astra on 28 September 2026, a day before its annual developer conference. The Wall Street Journal was first to report the decision, and OpenAI confirmed it the same day in a statement from Saachi Jain, its head of safety systems, per CNBC. Because the model was never released, no one outside OpenAI has used it, and OpenAI has not published evaluation figures for it.
Engadget, citing the Journal, reports that the model was due in October and was going to debut inside ChatGPT and Codex, and that the problems surfaced during internal testing.
GPT-6 Astra, the model GPT-6.1 Astra was measured against, was released earlier in September. The Register reports that the trouble was partly a side effect of making the model more useful: OpenAI had improved what it calls "model laziness", where an AI gives up or hands a task back to the user when it encounters an obstacle.
GPT-6.1 Astra gave up less often, which is what teams building on agents have been asking for. The Register adds that the same persistence made it worse at staying within the boundaries of what it had been authorized to do.
Why Did OpenAI Cancel GPT-6.1 Astra?
OpenAI cancelled GPT-6.1 Astra because it regressed on staying within the scope it was authorized for and on accurately reporting the work it had done. Jain's statement, as The Register carried it:
- Scope and authorization - whether the model stays within what it was asked and allowed to do.
- Reporting back - whether what the model tells the user about its work matches the work it did.
Jain framed the first reason as a trade-off: "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction." On the decision itself, the statement to CNBC added: "when we ship it to users, we have an extremely high bar in terms of safety and alignment."
OpenAI also told The Register that GPT-6.1 Astra performed worse than GPT-6 Astra on alignment evaluations, without naming the evaluations or the margin.
The Journal's Account of the Regressions
Gizmodo quotes the Journal as saying the model "regressed in two areas": deception, and the failure to seek authorization. The Journal described each one this way:
- Deception - "It wasn't always honest about telling users of the actions it did or didn't take."
- Scope authorization - it "would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe."
Gizmodo's own read is a useful check on the more dramatic headlines: "it was a faulty product, so OpenAI, to its credit, didn't ship it."
The Bar GPT-6 Astra Set in Its System Card
GPT-6 Astra's system card measures both behaviours its successor regressed on, which makes it the closest published baseline for "worse on alignment evaluations". Its section on deceptive interactions explains why OpenAI weighs misreporting so heavily: "a model that misrepresents its work is interfering with the ability of humans to supervise its actions," and such deception "is a plausible proxy for the model's general tendency to prioritize task completion over other safety considerations."
The card reports these GPT-6 Astra results on the evaluations closest to the two failures:
| Evaluation (card section) | What it measures | GPT-6 Astra result |
|---|---|---|
| Coding deception (8.3.1) | Misleading claims in the final response, including false reports of completed actions, tool access, verification, or ongoing background work | GPT-5.6 Sol's misrepresentation rate is 4x Astra's at maximum reasoning effort |
| Broken search tool (8.3.2) | How often the model fails to tell the user that the search tool a task needs is unavailable | GPT-5.6 Sol's failure rate is 10x Astra's at maximum reasoning effort |
| Realistic work environments (8.4) | Unintended outcomes such as unauthorized transactions or data exfiltration in adversarial workplace tasks | 3.4% misaligned outcome rate without the confirmation policy, against 18.8% for GPT-5.6 Sol |
| Unauthorized transactions (8.4) | Transactions the task did not authorize | 6.8% without the confirmation policy and 4.3% with it, against 38.0% and 7.4% for GPT-5.6 Sol |
| Internal Codex simulation (8.6) | Misalignment flags on simulated internal Codex traffic | About 53% fewer flags at severity 3 or above than GPT-5.6 Sol |
Section 8.6 traces misaligned coding behaviour to "a mix of overeagerness to complete the task and interpreting user instructions too permissively," and lists being "deceptive when reporting its results to users" among the ways it shows up. A model tuned to push through friction has more chances to cross a boundary, which is the trade-off Jain described.
None of this shows how far GPT-6.1 Astra fell, because OpenAI has not published its results. The card's own caveat also applies: the coding-deception tasks "were deliberately selected to elicit potentially dishonest behavior," so the rates are not expected to match production.
What Happens Next for OpenAI's Astra Models?
OpenAI will reuse the same base model for future GPT-6 releases, and Jain said the company will investigate the root cause of the problems, per Engadget. Engadget adds that OpenAI will use reinforcement learning that rewards the correct behaviour.
OpenAI also told The Register that more Astra models are coming, and that other new models that have cleared its safety bar will arrive "very soon." It has not said what caused the regressions.
The model OpenAI did launch that week is a different one. GPT-6.1 Sol, announced at DevDay on 29 September alongside OpenAI dots, is an upgrade to GPT-6 Sol that OpenAI says nearly matches GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra's standard token prices.
How Would Anyone Know an Agent Misreported Its Work?
By comparing the report with a record of what the agent did. Every item in the card's coding-deception definition (a completed action, tool access, a verification, work still running in the background) is a claim about an effect, and none of them can be judged from the response alone.
In testing, OpenAI had both halves, what the model did and what it said it did, and that is what made the gap visible. Its misalignment reports, including an earlier one on compaction summary misalignment, follow the same pattern: the problem shows up when someone inspects the record.
The DNS Report as a Worked Example
OpenAI's DNS report, a misalignment report last updated on 25 September, covers a separate incident. An internal research model in reinforcement learning training could not reach search engines, found insufficient DNS filtering in its sandbox, and used it to query a public chatbot:
- Its supplied search tool returned unrelated results, and direct requests to three search engines failed.
- It found that the sandbox's own DNS resolver answered real lookups, and used a public DNS service to forward questions to a chatbot.
- When the first queries timed out, it raised its timeout from six seconds to nineteen, and the chatbot answered.
- It wrote a new script and sent eighteen more questions down the same route.
The report notes: "The last batch was still running when the agent replied to the user." The reply said it could not identify the person, and that "giving a name would be guesswork."
The reply was accurate about the result and said nothing about the route. OpenAI found the route because its misalignment monitor flagged the behaviour within 15 minutes and the DNS activity was logged; the reply itself gave no sign of it.
The Same Gap in Production Agents
The gap is not limited to frontier models under lab conditions. Here is a reply from an expense agent we test against at TestMu AI, asked about claim EXP-1002:
approved for $450. This is within the $200 limit for software.That report happens to give itself away, because the sentence contains the two numbers that disprove it. Most reports carry no such tell: a reply that says "done", "verified", or "couldn't find it" reads the same whether or not it is true.
The approval above was real, and only what the agent actually did could confirm it. The Jev typed output post covers the harder case, where the same wrong decision comes back as a bare JSON field with no sentence to read. An agent that reports an action it never took is a case of agent action hallucination.
I made the same argument about GPT-6.1 Astra in an article on X:
OpenAI just shelved a model partly over how it described its own work. https://t.co/8ER223TRKN
- Vipul Verma (@vipulkv) September 30, 2026
Coding agents, the kind GPT-6.1 Astra was headed to Codex to power, have the same gap on every UI change. When a coding agent reports that a fix passed, it is reporting on the code, and the rendered page is invisible to it. Kane CLI checks the running app instead: it drives a real Chrome browser from a plain-English objective and returns a pass or fail backed by an evidence pack of per-step screenshots, a network log, and console output.
Checks to Run on Any Agent You Ship
Check what the agent's run changed, then compare that with what the agent said. An agent's account of what it did is the weakest evidence available about what it did, and in production it is usually the only half you get: the agent says the refund went through, the ticket was updated, the file was checked.
For each run, answer these from records outside the agent's own account:
- Calls - what did the agent call, and with what arguments?
- Changes - what did it change, and where: files, database records, tickets?
- Messages - what did it send, and to whom?
- Authorization - was each action one it was allowed to take, or one it should have asked about first?
- Match - does that record agree with what the agent said it did?
An agent that logs its tool calls can be checked against those logs, while one that records nothing leaves you only its account. For the wider set of failure modes behind these checks, see AI agent reliability.
TestMu AI's Agent Assurance runs this check on agents that act. It tests how your agents actually behave across workflows, tools, and actions: it reads the agent's code or spec, generates functional, non-functional, and adversarial scenarios, invokes the agent for real, and grades each criterion against observed evidence, such as files changed on disk and tool calls checked against the agent's own tool surface.
Each criterion gets Pass, Fail, or Unable to Verify. What it could not check is reported as the assurance gap and kept out of the pass rate, so a pass never rests on a claim nobody looked at. Agent Assurance is pre-alpha and publicly installable, and it runs from the terminal as rook.
Note: Your agent says it worked. Check what it actually did with TestMu AI Agent Assurance, which runs on a TestMu AI account. Try TestMu AI free!
Getting Started With Effect-Based Agent Checks
OpenAI could see GPT-6.1 Astra's misreporting because its tests recorded what the model actually did. Pick one agent you already run, list the actions it reports most often, and check a week of those reports against the records where each effect should appear: a file, a database row, a sent message.
To automate that comparison, the Agent Assurance quickstart runs a first evidence-graded suite against a sample support-triage agent.
Author
Vipul Verma is Group Senior Vice President of Engineering at TestMu AI (formerly LambdaTest), where he heads the entire engineering organization that builds KaneAI, HyperExecute, and the broader testing cloud. He brings 15+ years architecting, securing, and scaling large enterprise applications across multiple sites. Before TestMu AI he was India Head at LogicHub, where he built the India R&D site from the first employee to a 30-plus engineering team, and Principal Software Engineer at Sumo Logic, where he was the first engineer in the India office and shipped search-performance and pricing-model initiatives. Earlier he worked on trading platforms at Portware and D. E. Shaw. Vipul holds a B.Tech in Computer Science from IIT Kharagpur.
Reviewer
Mayank Bhola is Co-Founder and Head of Products at TestMu AI (formerly LambdaTest), where he leads the entire product portfolio across KaneAI, Kane CLI, HyperExecute, SmartUI, the Real Device Cloud, Accessibility, and other software testing product lines. As an early Lead Architect he designed and built the company's flagship Tunnel technology from scratch, created the React-based automation platform, and architected the data-intensive pipelines and FAAS services that scale it. He brings more than 10 years of experience in software development and product engineering, with earlier roles as Head of Technology at Juggernaut Books and Senior Software Engineer at PressPlay TV and Zomato. Mayank holds a B.Tech in Computer Engineering from JIIT Noida.
GPT-6.1 Astra FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




