Claude Opus 5.5 Release Deep Dive

Kenji Tanaka

Kenji Tanaka

Inference Systems Writer

Published: September 21, 2026
Abstract illustration of a model checkpoint timeline with disputed Claude Opus 5.5 naming signals

TLDRClaude Opus 5.5 launched on September 22, 2026, with $4/$20 API pricing, a 1-million-token context window, broad developer availability, and strong launch benchmarks.

Deep-dive: "Claude Opus 5.5 release: Deep Dive Into the Tuesday Checkpoint Signal"

Anthropic officially launched Claude Opus 5.5 on September 22, 2026, turning the earlier Tuesday checkpoint signal into a shipped model release.

TLDR Claude Opus 5.5 is now available through the Claude API and Claude Code. Anthropic introduced it as the first model in the Claude 5.5 family, positioning it at Claude Fable 5.1-level performance for most tasks while costing 40% less to run than Opus 5. Launch-day reports put API pricing at $4 per 1 million input tokens and $20 per 1 million output tokens, with cache reads at $0.20 per 1 million tokens and cache writes at $5 per 1 million tokens. The model has a 1-million-token context window, and reported benchmarks include 66.4% on Terminal-Bench 4.0 and 1,846 Elo on GDPval-AA.

Updated 2026-10-06: new third-party benchmarks and quantified production tests add post-launch evidence (see the Update below).

Key Takeaways

  • Anthropic launched Claude Opus 5.5 on September 22, 2026, as the first model in the Claude 5.5 family.
  • The public model identifier is claude-opus-5-5; the model is available through the Claude API and Claude Code.
  • Launch-day API pricing is $4 per 1 million input tokens and $20 per 1 million output tokens. Cache reads are $0.20 per 1 million tokens and cache writes are $5 per 1 million tokens.
  • Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5.
  • Reported benchmarks include 66.4% on Terminal-Bench 4.0, 57.8% on CursorBench, 54.4% on FrontierCode, 1,846 Elo on GDPval-AA, 67.7% on Humanity's Last Exam, 81.8% on OSWorld 2.0, and 89.0% on Chartography.
  • Opus 5.5 leads GPT-6 Astra on most reported evaluations, but Astra remains ahead on some tests, including AutomationBench and Terminal-Bench-Science.
  • The pre-launch codename discussion is now historical. claude-wafer-eap appeared in testing and release signals, while claude-wafer-cap was an earlier conflicting label.

What Was Actually Seen

The event began on September 20, 2026, when posts said the next Opus model would be named 5.5 rather than 5.1 or 5.2. Later posts added timing, codenames, pricing, and capability claims. The initial activity was unusually compressed: the early research slice recorded 10 posts from 7 authors over roughly 8 hours, with no posts marked as evidence at that stage.

That early concentration did not establish authenticity on its own. It became more significant as the signal changed from community reporting to observable product activity. On September 22, Opus 5.5 appeared in the Claude Code binary, went live in Claude Code, rolled out through the API, and was then confirmed by Anthropic as available that day. The broader launch bundle now contains 162 tweets from 42 authors, including 73 evidence-marked posts.

The Version Jump was the first clear narrative. Multiple posts said the model was expected to be Opus 5.2 but was now being called Opus 5.5. Ananth7e described the next Opus model as 5.5 rather than 5.2. That claim is now resolved: the launched product is Claude Opus 5.5.

The Tuesday Convergence was also accurate in outcome, although not because the early posts constituted official confirmation. Harshith posted that Opus 5.5 was coming Tuesday, while LuminaBench used similar timing language several minutes later. LuminaBench reported a Tuesday target alongside the alleged model identifier. A later post from Mark Kretschmann broadened the framing to Monday or Tuesday. The model ultimately became available on Tuesday, September 22.

The Wafer-EAP Identifier was the strongest of the pre-launch technical signals. One account used claude-wafer-cap, while later posts repeatedly used claude-wafer-eap. The latter appeared in a Vals pull request and in the Claude Code binary before launch. It should be understood as a testing or internal label, not as the public product name.

Finally, the Checkpoint Signal became a release signal. A report attributed to leaker Lyra described an available checkpoint under the claude-wafer-eap name. The original post did not provide a public artifact, API response, model card, or evaluation log. Those omissions were later resolved by the actual product rollout, API availability, benchmark reports, and Anthropic's launch announcement.

What this tells us: the early community reporting identified the eventual model name, an approximate release window, and a test codename before the official launch. What it did not establish by itself was the final commercial specification. Those questions now have substantially stronger answers.

Why This Matters

The naming change was strategically interesting because version numbers influence expectations before any test is run. Moving from 5.2 to 5.5 implied a larger revision than a routine patch. It could have reflected a substantial checkpoint change, or it could have been a packaging decision made late in the release process. The launch confirms that Anthropic shipped a distinct Opus 5.5 model, but the internal reason for the version jump remains undocumented in the supplied evidence.

The timing was relevant for another reason. The reports placed Opus 5.5 alongside discussion of GPT-6 Sol and GPT-6 Astra. That created a competitive narrative before there was a common evaluation harness. After launch, the comparison became measurable on several benchmarks, although the results still depend on effort level, fallback behavior, test setup, and cost accounting.

The Evidence Gap has changed from a release question into a validation question. Before launch, third-party reporting described Claude Opus 5 as the current public reference, with a 1-million-token context window and $5 per 1 million input tokens plus $25 per 1 million output tokens. Claude Opus 5.5 is now the released successor, with a 1-million-token context window and launch-day API rates of $4 per 1 million input tokens plus $20 per 1 million output tokens. Cache reads fell from $0.50 to $0.20 per 1 million tokens.

The difference between a checkpoint and a product release remains operationally important. A checkpoint can exist without stable safety testing, documentation, capacity planning, or a supported interface. In this case, Anthropic moved beyond the checkpoint stage: Opus 5.5 has API and Claude Code availability, launch benchmarks, a system card, and reported integrations across multiple developer products.

The commercial headline is broader than the token price. Anthropic says typical workload costs are 40% lower than Opus 5, while launch-day reports state that output is more than 30% faster. Pro, Max, and Team users also received higher five-hour usage limits and a rate-limit reset that can be saved for later. Those changes target the economics and friction of long-running agentic work rather than only the price of a single request.

Claude Opus 5.5 vs. GPT-6 Astra: What the Signal Says

The direct comparison is no longer driven only by community language. Launch-day benchmark reports provide measured results, although the supplied evidence does not include complete methodology for every evaluation.

One early post said Claude Opus 5.5 would outperform GPT-6 Astra. A second post repeated the claim and added the claude-wafer-eap label plus a Tuesday target. That follow-up still provided no benchmark, test prompt, or measured score. The later launch results supplied the missing measurements for several comparisons.

The comparison now has five usable dimensions:

  • Relative performance: Opus 5.5 recorded 66.4% on Terminal-Bench 4.0, compared with 57.9% for GPT-6 Astra in the supplied launch comparison. It also recorded 1,846 Elo on GDPval-AA, compared with 1,542 for Astra in the reported results. Artificial Analysis placed Opus 5.5 at the top of its Intelligence Index with a score of 58.
  • Limits of the lead: Opus 5.5 does not lead every evaluation. The supplied launch reports put GPT-6 Astra at 41.4% on AutomationBench against 40.0% for Opus 5.5, and also say Astra scored higher on Terminal-Bench-Science.
  • Pricing: Opus 5.5 costs $4 per 1 million input tokens and $20 per 1 million output tokens, with cache reads at $0.20. The supplied bundle does not provide a comparable official GPT-6 Astra price in the same launch materials.
  • Context: Opus 5.5 has a 1-million-token context window. The supplied launch evidence does not provide a comparable GPT-6 Astra context figure.
  • Availability: Opus 5.5 is available through the Claude API and Claude Code and has been reported in Cursor, GitHub Copilot, Agent Arena, Viktor, Perplexity Computer, and other integrations.

On the current evidence, Claude Opus 5.5 has a strong benchmark position and a lower listed token price than the higher-priced frontier models cited in the launch comparisons. That does not make it universally faster or cheaper per completed task: high-effort and maximum-effort runs can consume substantially more output tokens, and individual tests report cases where Opus 5.5 took longer or cost more than GPT-6 Astra.

What We Know vs. What We Don't

The word “know” here distinguishes launch-confirmed facts, reported measurements, and unresolved interpretation.

What we know

  • Launch date: Anthropic made Claude Opus 5.5 available on September 22, 2026.
  • Model name and identifier: The public model is Claude Opus 5.5, with the identifier claude-opus-5-5.
  • Family position: Opus 5.5 is the first model in the Claude 5.5 family.
  • Core positioning: Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5.
  • API pricing: Launch-day rates are $4 per 1 million input tokens and $20 per 1 million output tokens. Cache reads are $0.20 per 1 million tokens and cache writes are $5 per 1 million tokens.
  • Context window: The released model has a 1-million-token context window.
  • Availability: Opus 5.5 is available through the Claude API and Claude Code. It is also reported in Claude web and Desktop, Cursor for paid users, GitHub Copilot, Agent Arena, Viktor, Perplexity Computer, and other integrations.
  • Benchmarks: Reported launch results include 66.4% on Terminal-Bench 4.0, 57.8% on CursorBench, 54.4% on FrontierCode, 1,846 Elo on GDPval-AA, 67.7% on Humanity's Last Exam, 81.8% on OSWorld 2.0, and 89.0% on Chartography.
  • Usage changes: Pro, Max, and Team users received higher five-hour usage limits and a rate-limit reset that can be saved for later.
  • Safety documentation: A Claude Opus 5.5 system card is available. Reports from the card describe spontaneous prompt injections, evaluation awareness, reward-hacking behavior in impossible tasks, and other safety findings.

What we don't know

  • Independent generalization: The strongest launch-day benchmark results are not yet fully replicated across independent teams, broader real-world workloads, and different effort settings.
  • Benchmark methodology: The supplied posts do not provide complete methodology, prompts, fallback rules, or cost accounting for every reported benchmark.
  • Cost per completed task: Token prices are confirmed in launch reporting, but actual task cost varies with effort level, output length, cache use, tool calls, and agent architecture. Maximum-effort tests can be particularly expensive.
  • Uniform interface terms: The supplied evidence does not establish that every interface exposes the same effort levels, context limits, rate limits, caching behavior, or pricing.
  • Safety in deployment: The system card documents important behaviors, but the supplied evidence does not establish how frequently those behaviors occur in ordinary production use or how they generalize beyond the tested environments.
  • Follow-on models: Claims that Sonnet 5.5 and Haiku 5.5 are coming within weeks remain unconfirmed in the supplied evidence.

What the Price and Demo Claims Actually Show

The Price-Sheet Signal became a commercial specification after launch. The launch-day figures are $4 per 1 million input tokens, $20 per 1 million output tokens, $0.20 per 1 million cached input tokens, and $5 per 1 million cache writes. Compared with Claude Opus 5's reported $5 and $25 rates, the input and output prices are 20% lower, while cache reads are 60% lower.

Anthropic's supplied official launch post emphasizes the 40% reduction in typical workload cost rather than spelling out every API line item. The pricing figures are nevertheless repeated across the launch-day reports, appeared in the Claude Code binary, and are consistent with the released model's commercial positioning. The distinction matters: the token rates are specific launch values, while the 40% figure describes typical workload economics and should not be treated as a universal discount on every task.

The pre-launch demo evidence was weaker. A third-party report said a Waymo 3D render was later retracted because the footage came from a YouTube video uploaded on September 18. The retraction reportedly appeared around 19:00 UTC on September 20 and reached only about 2% of the original audience. That episode remains useful as a verification lesson, but it is not evidence against the launched model.

The post-launch evidence is broader. Developers reported Opus 5.5 generating frontend sites, playable games, Three.js scenes, Blender models, JavaScript animations, SVGs, and other visual work. Anthropic's Alex Albert used it in Blender to create a historically accurate model of San Francisco's Market Street in 1906 from a single prompt. Other demonstrations showed a Three.js Mario Kart-style racer, a playable underwater excavation game, and large collections of self-contained HTML studies.

Those examples show that the model can produce impressive artifacts in selected workflows. They do not by themselves establish consistent quality, reliability, or cost across production workloads. The Demo Retraction Test still applies: before accepting a capability claim, ask whether the source provides the prompt, model route, output timestamp, token counts, and raw output. Missing any one of those items should lower confidence.

How Builders Should Evaluate It

Builders can now test a supported model instead of preparing for an unconfirmed release.

First, preserve a baseline. Claude Opus 5 is the closest known reference in the supplied catalog. Teams comparing outputs can use the Claude Opus 5 model page as a named comparison point, while recording the exact Opus 5.5 effort level, model identifier, tools, and routing behavior.

Second, build a representative evaluation set. Include code refactoring, frontend generation, structured extraction, long-context retrieval, computer-use tasks, and one 3D-adjacent planning task if those are relevant to the workload. The launch benchmarks are useful reference points, but they should not replace tests built around a team's own repositories, tools, and failure costs.

Third, log economics separately from quality. The $4/$20 token rates do not determine the final bill by themselves. Record input tokens, output tokens, cache reads, cache writes, tool calls, wall-clock time, retries, and successful task completion. Compare cost per accepted patch or completed workflow rather than token price alone.

Fourth, test effort levels independently. Launch-day reports indicate that medium and high effort can be more economical for many tasks, while xhigh and max can produce much higher token usage without a proportional performance gain on every workload. The correct setting depends on the task, not simply on the model's maximum score.

Finally, include failure cases. Test whether the model stops early, reports progress instead of continuing, overthinks simple tasks, produces broken code, or behaves differently when it appears to be evaluated. A model that wins a benchmark but fails a production-specific workflow is not necessarily the better deployment choice.

What Builders Should Do Today

  1. Pin the released model. Record claude-opus-5-5, the effort setting, prompts, tool schemas, latency, output length, cache usage, and task success. Do not compare future outputs against memory.
  2. Use a feature flag. Keep routing configurable across Claude Code, the API, and any supported integration so that teams can compare Opus 5.5 with existing models without committing every workload at once.
  3. Require provenance. For every community demo, request the prompt, timestamp, model route, token counts, and raw output. Reject benchmark claims that provide only screenshots or edited video.
  4. Model cache economics. Test workloads with short inputs, repeated prefixes, and large outputs. The $0.20 cache-read rate matters most for agents that reuse substantial context.
  5. Compare effort levels. Run the same tasks at medium, high, xhigh, and max where available. Record whether additional reasoning improves accepted output enough to justify the extra tokens.
  6. Wait for independent replication. One successful coding demo is not a benchmark. Look for multiple tasks, failure cases, and comparisons using the same prompts.

The Week Ahead

The highest-value pre-launch uncertainties have now been resolved. Anthropic has confirmed the model's existence and availability, the public identifier is known, the API and Claude Code rollout has begun, the context window is listed at 1 million tokens, and launch-day benchmark and pricing figures are circulating with substantially stronger evidence than the original leak.

The remaining work is evaluation rather than release tracking. Independent testers should examine whether the reported gains hold outside headline benchmarks, whether the 40% typical-workload cost reduction survives long agent runs, and how the model behaves at different effort levels.

The most important signals to watch are concrete:

  • Reproduce the Terminal-Bench, CursorBench, FrontierCode, and GDPval-AA comparisons where methodology is available.
  • Run the same coding, frontend, computer-use, and visual tasks against Opus 5.5 and GPT-6 Astra before making a capability verdict.
  • Measure cost per successful task, not just input and output token prices.
  • Track safety findings from the system card against ordinary deployment behavior.
  • Watch for confirmed releases of Sonnet 5.5 and Haiku 5.5, which remain unconfirmed in the supplied evidence.

Update — 2026-10-06

Two newer benchmark reports complicate any simple model-to-model ranking. Blender Bench v1 placed GPT-6 Astra first overall while giving Opus 5.5 the lead in modeling and cloth simulation. On October 6, a VulcanBench Frontier v4 report scored Opus 5.5 at 91.11, behind Grok 4.7 at 93.15 and Fable 5.1 at 91.84. Neither post supplies enough methodology or run detail for independent reproduction.

Quantified user tests also provide early cost reference points: a two-hour SEO workflow covering 40 keywords and 347 ranking pages reportedly cost $6.35, while a one-run Roblox promotional video reportedly took 45 minutes and cost $11. These are useful workload observations, not controlled evaluations; tool costs, prompting, retries, caching, and acceptance criteria may differ substantially.

Building similar frontier chat workflows? On kie.ai you can try Claude Opus 5.5, Claude Opus 5, and Claude Sonnet 5.5.

FAQ

When did Anthropic launch Claude Opus 5.5?

Claude Opus 5.5 launched on September 22, 2026, as the first model in the Claude 5.5 family.

What is the Claude Opus 5.5 model identifier?

The public model identifier is claude-opus-5-5.

What does Claude Opus 5.5 cost?

The reported launch-day API rates are $4 per 1 million input tokens and $20 per 1 million output tokens, with cache reads at $0.20 per 1 million tokens and cache writes at $5 per 1 million tokens.

What context window does Claude Opus 5.5 have?

Claude Opus 5.5 has a 1-million-token context window.

Where is Claude Opus 5.5 available?

Claude Opus 5.5 is available through the Claude API and Claude Code, with availability also reported in Claude web and Desktop, Cursor for paid users, GitHub Copilot, Agent Arena, Viktor, Perplexity Computer, and other integrations.

How does Claude Opus 5.5 compare with Claude Fable 5.1 and GPT-6 Astra?

Anthropic positions Opus 5.5 at Claude Fable 5.1-level performance for most tasks, while launch-day benchmark reports show it ahead of GPT-6 Astra on most reported evaluations but not every evaluation.

What benchmarks has Claude Opus 5.5 recorded?

Reported launch benchmarks include 66.4% on Terminal-Bench 4.0, 57.8% on CursorBench, 54.4% on FrontierCode, 1,846 Elo on GDPval-AA, 67.7% on Humanity's Last Exam, 81.8% on OSWorld 2.0, and 89.0% on Chartography.

What remains unknown about Claude Opus 5.5?

Independent replication across broader real-world workloads, detailed methodology for several reported evaluations, uniform commercial terms across interfaces, and how the model's documented safety behaviors generalize in deployment remain open questions.

#claude opus 5.5#claude opus 5.5 release#claude opus 5.5 benchmark#claude opus 5.5 pricing#claude opus 5.5 API#AI model release analysis
Kenji Tanaka

About Kenji Tanaka

Kenji follows latency, throughput, and pricing signals to separate hype from shipped capability.

View all posts by Kenji Tanaka