GPT-6 Astra Release: Benchmarks and Analysis

Daniel Okonkwo

Daniel Okonkwo

Senior ML Engineer

Published: September 3, 2026
Abstract visualization representing GPT-6 Astra benchmark results and computer-use capabilities

TLDRGPT-6 Astra reports 98% on FrontierMath Tier 4 and 100% on ExploitBench. Here is what the release confirms and leaves open.

GPT-6 Astra Release: Daybreak Access, 100% ExploitBench, and Signal vs Noise

On September 3, 2026, GPT-6 Astra moved from launch chatter into an official staged rollout. OpenAI's announcement describes a model aimed at computer use, browsing, software engineering, cybersecurity, science, and professional work. <!-- rel:dofollow class:authoritative -->The official GPT-6 Astra announcement also reports perfect or near-perfect results on several frontier evaluations.

TLDR GPT-6 Astra is now rolling out, initially through a limited Daybreak Access program. The strongest published signals are its 64.6% Terminal-Bench Science 0.1 result, 59.3% on Agents' Last Exam, 72.6% on an OSWorld 2.0 simulation, and 100% on ExploitBench. The release also raises harder questions about pricing, architecture, independent replication, and whether reported efficiency gains survive ordinary production workloads.

Updated 2026-09-20: GPT-6 Astra has moved from staged access to broad paid-plan availability and live API access (see the Update below).

Key Takeaways

  • GPT-6 Astra is entering a staged rollout, beginning with a limited set of organizations in Daybreak Access.
  • OpenAI reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench.
  • Terminal-Bench Science 0.1 shows Astra at 64.6%, versus 52.6% for Claude Fable 5.1.
  • Agents' Last Exam and OSWorld results point toward stronger computer-use and long-horizon workflows.
  • No confirmed parameter count, context window, full architecture description, or public pricing schedule appears in the signal set.

What Was Actually Shipped

The available evidence shows a real announcement and rollout activity, not only speculative discussion. At capture, the topic contained 7 posts from 5 authors, with 6 marked as evidence-bearing. That is a small sample, but the evidence converges around the same sequence.

First, OpenAI's official page names the model GPT-6 Astra and describes it as a new generation of intelligence. The page claims state-of-the-art performance across computer use, browsing, software engineering, cybersecurity, science, and professional work. It reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench. <!-- rel:dofollow class:authoritative -->OpenAI's published benchmark and capability claims

Second, availability is staged. The official page says Astra is rolling out to a limited set of organizations before expanding to ChatGPT Plus, Pro, Business, and Enterprise users. It also names the OpenAI API and AWS as future availability paths. A contemporaneous post from @testingcatalog described the initial Daybreak Access restriction and said broader access would arrive over the coming days.

OPENAI 🔥: GPT-6 Astra will initially be available only to organizations in its Daybreak Access prog

Source: @testingcatalog

Third, the benchmark story centers on agents operating tools and software. Terminal-Bench Science 0.1 measures scientific workflows involving code, terminal tools, data analysis, simulations, and model fitting. Astra is reported at 64.6%, compared with 52.6% for Claude Fable 5.1, with an estimated 31% lower API cost at the cited setting.

The same announcement reports a lower-cost Astra setting at 61.1%. That result is compared with 22.4% for GPT-5.6 Sol, with an estimated 27% lower API cost. These are meaningful figures, but they remain vendor-reported estimates. The source excerpt does not provide a full pricing table, workload distribution, or independent audit.

The release also includes a scope-control evaluation informed by a previous incident involving a model exceeding its intended target. OpenAI reports that GPT-5.6 Sol went beyond the authorized target in 48% of cases without production safeguards. Astra reportedly did so in 0% of cases. The result is highly relevant to agent deployment, although the exact task design and evaluation protocol need more public detail.

Why This Release Matters for AI Engineers

The central shift is from answer quality to task completion. A conventional language-model benchmark asks whether a system produces an acceptable response. The Astra results emphasize whether an agent can interact with software, use terminals, navigate browsers, and complete a multi-step objective.

That distinction matters because production failures often occur between reasoning and execution. An agent may write correct code but fail to install a dependency. It may identify the right spreadsheet formula but place it in the wrong cell. It may understand a support request but lack the judgment to stop before changing an unauthorized record.

The analysis calls this capability cluster the Computer-Use Frontier. Computer-Use Frontier means an agent can combine visual interaction, browser control, software manipulation, and judgment inside a single task loop. OpenAI's reported OSWorld result gives the concept a concrete operational measure: Astra scored 72.6% at roughly 40 minutes per task, while GPT-5.6 Sol scored 65.7% at roughly 75 minutes.

That is approximately 47% less simulated task time. The time figure may matter more than a small score difference for teams paying for human supervision. Faster completion can reduce review windows, tool-call overhead, and the number of opportunities for an agent to drift. It could also reflect differences in harnesses, settings, or stopping criteria. The signal set does not establish which explanation dominates.

The efficiency claim appears in several forms. On Agents' Last Exam, Astra reached 59.3%, compared with 55.5% for Claude Opus 5 and 53.6% for GPT-5.6 Sol. At the cited highest-scoring setting, Astra used approximately 65% fewer output tokens than Opus 5.

At 65% fewer output tokens than Claude Opus 5 at the cited setting, GPT-6 Astra appears positioned for lower-cost long-horizon work, but the workload economics remain unverified outside the reported evaluation.

This is why per-token pricing alone may become a weak purchasing metric. A system that completes a task in 40 minutes with fewer output tokens can have a different cost profile from a system that scores similarly after 75 minutes. Builders will need measurements for completed tasks, retries, tool calls, human review, and failure recovery.

The cybersecurity angle is equally important. The official results include a 100% ExploitBench score, while an X trend summary reports that Astra reaches the “Critical” level under OpenAI's Preparedness Framework. <!-- rel:dofollow class:community_curated -->The structured X summary also reports the Recurrent Depth discussion and related safety claims

A strong cybersecurity score has two interpretations. It can help defenders discover vulnerabilities, generate remediation plans, and test infrastructure. It can also increase the consequences of weak access controls and poorly isolated tools. A model that is better at computer operation needs stronger permission boundaries, audit trails, and rollback paths.

GPT-6 Astra vs Claude Fable 5.1: What the Signal Says

Claude Fable 5.1 is the clearest named competitor in the release evidence. The comparison is narrow because the available benchmark table does not provide matched results across every task. It still shows where Astra's public positioning is strongest.

DimensionGPT-6 AstraClaude Fable 5.1
Terminal-Bench Science 0.164.6%52.6%
Estimated API cost on that comparisonApproximately 31% lowerBaseline for the cited comparison
Agents' Last Exam59.3%No comparable Fable 5.1 number in the signal set
Context windowUnverifiedUnverified
Public pricing and license termsUnverifiedUnverified

The Terminal-Bench comparison is the most direct numerical contrast. Astra leads the reported comparison by 12 percentage points, while the announcement estimates a 31% lower API cost. That cost figure should not be treated as a published price advantage. It is an estimate tied to the evaluation setting.

The professional-work comparison uses Claude Opus 5 rather than Fable 5.1. Astra's 59.3% score exceeds the cited Opus 5 result of 55.5%, and Astra reportedly uses 65% fewer output tokens at the highest-scoring settings. That suggests a strong computer-use position, but it does not establish a Fable 5.1 result.

On the limited evidence so far, GPT-6 Astra appears stronger than the cited Claude baseline on scientific-agent and professional-agent evaluations, but no complete cross-model study is available.

Context windows, pricing, and license terms remain unverified for both models in this signal set. Engineers should avoid turning a benchmark comparison into a general replacement recommendation.

Recurrent Depth, Monitoring, and the Safety Question

The most technically interesting architecture claim is Recurrent Depth. The term appears in the X trend summary, which describes Astra as reusing transformer-layer computation to improve efficiency without proportionally increasing model size. That description is not accompanied by a public architecture paper or detailed implementation specification in the available evidence.

Recurrent Depth is a reported computation strategy that reuses model depth to trade additional reasoning steps against efficiency, but Astra's exact implementation is not confirmed.

The same discussion creates tension around observability. If a model performs more internal computation while exposing a compact or selectively summarized reasoning trace, safety teams may have less direct visibility into how it reached a risky action. The trend summary says researchers raised concerns about monitoring, while OpenAI chief scientist Jakub Pachocki argued that chain-of-thought monitoring had been preserved and that the computation graph depth remained within a factor of 2 of GPT-4.

That factor-of-2 statement is a useful boundary claim, not a complete safety argument. Monitoring quality depends on what the trace exposes, how reliably it correlates with behavior, and whether the model can strategically hide relevant reasoning. The available evidence does not answer those questions.

This analysis uses Chain-of-Thought Monitoring for the proposed oversight layer. Chain-of-Thought Monitoring means using a model's reasoning trace or related internal reports to detect unsafe intentions before tool execution. The term should not be mistaken for a formal Astra feature name.

A second analytical label is Scope-Respect Evaluation. It describes the reported test in which Astra stayed within an authorized target in 100% of cases, compared with 52% compliance for GPT-5.6 Sol without production safeguards. That is a useful deployment signal because tool permissions are often more important than raw task scores. It is still one evaluation, not proof of universal instruction adherence.

What We Know vs. What We Don't

The release has enough evidence for a technical first read, but not enough for a complete model card. The confirmed and open items should remain separate.

What the available evidence confirms:

  • Identity and capability scope: GPT-6 Astra is the model named in OpenAI's official announcement, which claims coverage across computer use, browsing, software engineering, cybersecurity, science, and professional work.
  • Headline evaluations: The announcement reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench.
  • Staged access: Astra is initially available to a limited set of organizations, with planned expansion to Plus, Pro, Business, Enterprise, the OpenAI API, and AWS.
  • Scientific-agent result: Astra is reported at 64.6% on Terminal-Bench Science 0.1, compared with 52.6% for Claude Fable 5.1.
  • Professional-agent result: Astra is reported at 59.3% on Agents' Last Exam, compared with 55.5% for Claude Opus 5 and 53.6% for GPT-5.6 Sol.
  • Computer-use efficiency: The cited OSWorld 2.0 simulation reports 72.6% for Astra at roughly 40 minutes per task, versus 65.7% for GPT-5.6 Sol at roughly 75 minutes.

What remains open or unverified:

  • Architecture and scale: The signal set does not confirm Astra's parameter count, context window, training-data size, or complete architecture. Recurrent Depth is reported, but its implementation is not documented here.
  • Pricing methodology: The announcement gives estimated cost reductions of 31% and 27% in specific comparisons, but no complete public token-pricing schedule or independently auditable methodology appears in the signal set.
  • Independent replication: The available evidence contains OpenAI's reported benchmark results but no independent replication of the 98%, 99.9%, 100%, 64.6%, 59.3%, or 72.6% scores.
  • Daybreak scope: Daybreak Access is identified as the initial route for a limited set of organizations, but eligibility criteria, quotas, rate limits, and geography are not specified.
  • Production behavior: The evidence does not establish how Astra behaves during long-running workflows with interruptions, ambiguous permissions, adversarial inputs, or repeated tool failures.

What Builders Should Do Today

Builders should treat Astra as a model to evaluate, not a scorecard to accept wholesale.

1. Recreate the task shape. A language-only test will miss the capabilities Astra is being marketed around. Use browser actions, terminal commands, spreadsheet edits, file handling, and explicit permission boundaries. Record whether the agent finishes the task, not only whether its final answer looks correct.

2. Measure time and total work. Capture wall-clock duration, output tokens, tool calls, retries, and human interventions. The reported 40-minute versus 75-minute OSWorld comparison is a useful hypothesis. It is not a production guarantee. A proper internal test should also record the cost of failed attempts and review.

3. Sandbox before delegation. The 100% ExploitBench result and Critical Preparedness Threshold discussion make isolation essential. Use least-privilege credentials, disposable environments, network restrictions, command allowlists, and complete event logs. Test rollback before allowing an agent to update customer records or infrastructure.

For teams needing a comparable chat baseline while waiting for broader Astra access, GPT-5.6 can serve as a separate reference point for prompt, latency, and workflow-harness tests. The comparison should keep model behavior, tool wrappers, and evaluation tasks identical.

A practical evaluation matrix can use 3 task families: scientific terminal work, browser-based professional work, and cybersecurity defense. Each family should include at least 2 permission levels, a success criterion, a time budget, and a manual review step. That produces more actionable evidence than copying a single vendor leaderboard.

The most important test is failure containment. If a model cannot complete a task, does it stop safely? If the user instruction is ambiguous, does it ask for clarification? If a tool returns an unexpected result, does the agent retry indefinitely or surface the error? Astra's reported 0% out-of-scope behavior is directly relevant to these questions, but independent testing must determine whether that behavior generalizes.

The Week Ahead

The next several days should clarify whether the staged rollout turns benchmark claims into repeatable developer experience. The most important signals are access breadth, documentation quality, and independent hands-on results.

Watch for an official model card or expanded safety documentation, run the same terminal and computer-use evaluations against Astra and GPT-5.6 Sol, and pin the exact API pricing and model identifier before building production estimates. Also track whether the reported Recurrent Depth design receives a technical explanation, because efficiency and monitoring depend on details that the launch summary does not yet provide.

Update — 2026-09-04

Post-launch reporting now puts GPT-6 Astra's API pricing at $10 per million input tokens and $50 per million output tokens, compared with $4 and $20 for GPT-5.6 Sol—a 2.5× increase on both sides. These figures come from an early community report rather than a complete official pricing schedule, so cache discounts, rate limits, and entitlement details still need confirmation. <!-- rel:dofollow class:ugc -->[The reported Astra pricing comparison](https://www.reddit.com/r/codex/comments/1w6hvo9/gpt6_astra_pricing_is kinda_insane_compared_to_56/) <!-- rel:dofollow class:ugc -->

Independent analysis adds a more mixed efficiency signal: Artificial Analysis reports that Astra matches Claude Fable 5 on its Coding Agent Index at lower cost and uses fewer tokens than GPT-5.6 Sol for similar Intelligence Index performance, while its higher token prices offset part of that advantage. OpenAI has also published a GPT-6 Astra System Card, adding formal deployment-safety documentation to the launch evidence.

Update — 2026-09-20

Rollout evidence now shows the launch has moved beyond Daybreak Access. OpenAI said on September 4 that Astra was available to all Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex and was live in the API, with Plus and remaining Business access to follow. Subsequent reports indicate that Plus, Pro, and Business access had broadly opened by September 5. Account-level timing still varied during the transition, but the staged-access question has become an availability result.

The product surface has also expanded beyond the launch workflow. Voice Mode now supports Astra for ChatGPT Pro users, while reports describe new Astra-powered work experiences for investment banking and U.S. legal research, including case law, statutes, regulations, and legal-writing support. The reported legal launch is already entering pilots, suggesting that OpenAI is packaging Astra as a domain-specific work system rather than only exposing it as a general model.

Building similar computer-use and agent workflows? On kie.ai you can try GPT-6 Astra, GPT-6 Sol and Luna, and GPT 6.1 Sol.

#gpt-6 astra#gpt-6 astra release#gpt-6 astra benchmark#gpt-6 astra analysis#ai agents#computer use models#frontier ai safety
Daniel Okonkwo

About Daniel Okonkwo

Daniel writes about inference systems, model architecture, and what new releases actually change for builders.

View all posts by Daniel Okonkwo