GPT-5.6 Deep Dive: Sol, Terra, Luna Release

Daniel Okonkwo

Daniel Okonkwo

Senior ML Engineer

Published: July 10, 2026
GPT-5.6 Sol, Terra, and Luna release banner with three-tier model illustration

TLDROpenAI shipped GPT-5.6 as Sol, Terra, and Luna with ultra mode and multi-agent APIs. What's actually in the release, and what to watch.

GPT-5.6 Deep Dive: Sol, Terra, and Luna Land With Ultra Mode and Multi-Agent APIs

At 17:02 UTC on July 9, 2026, Sam Altman confirmed the GPT-5.6 livestream was live, and within six minutes the model was rolling out globally. No staged preview. No week-long US-first window. A single livestream, three model tiers, and a rebranded desktop app — all shipped in one afternoon.

TLDR OpenAI released GPT-5.6 as a three-tier family — Sol (flagship), Terra (balanced), Luna (cheap) — on July 9, 2026. The family introduces a new Max reasoning setting, Ultra Mode for parallel agent coordination, Programmatic Tool Calling that lets the model write JavaScript to orchestrate tools, and a Multi-agent beta in the Responses API. Pricing lands at $1/$6, $2.50/$15, and $5/$30 per 1M input/output tokens. OpenAI claims Sol beats Claude Fable 5 by 13.1 points on Agents' Last Exam, but Fable 5 still leads Sol on SWE-Bench Pro (80% vs 64.6%). ChatGPT Work, a rebranded desktop app, and hosted Sites shipped alongside the model.

Key Takeaways

  • GPT-5.6 ships in three named tiers: Sol, Terra, Luna — replacing the older Pro/Mini naming scheme.
  • A new Max Reasoning Effort setting sits above the existing xhigh tier, and a separate Ultra Mode coordinates multiple subagents in parallel.
  • The API adds Programmatic Tool Calling (JS-based tool orchestration), a Multi-agent beta, and explicit prompt cache breakpoints.
  • Pricing per 1M input/output tokens is Luna $1/$6, Terra $2.50/$15, Sol $5/$30 — roughly half Claude Fable 5's headline rate.
  • Sol beats Claude Fable 5 by 13.1 points on OpenAI's own Agents' Last Exam number, but trails Fable 5 on SWE-Bench Pro (64.6% vs 80%).
  • The GPT-5.4 sunset is dated July 23, and the Atlas browser is scheduled to sunset around August 9.

What Actually Shipped Today

The official release page went up the same morning, and OpenAI Developers confirmed the three-tier split on the API. The concrete facts, drawn only from primary sources posted in the last few hours:

  • Three model tiers, generally available: Sol (flagship, long-horizon coding and agentic work), Terra (balanced everyday work), Luna (cheap, high-volume, latency-sensitive work).
  • API pricing per 1M input/output tokens: Luna $1/$6, Terra $2.50/$15, Sol $5/$30, per Simon Willison's same-day analysis.
  • Six reasoning effort levels available: none, low, medium, high, xhigh, and a new max — confirmed in the OpenAI model guidance docs.
  • Ultra Mode, described on the release page as the highest-capability setting, which "coordinates multiple agents across parallel workstreams."
  • Programmatic Tool Calling: the model writes JavaScript that orchestrates tool calls in a hosted runtime, per the developer guide.
  • Multi-agent [beta] in the Responses API: a GPT-5.6 instance can spin up subagents in parallel and synthesize their results.
  • Explicit prompt cache breakpoints, plus persisted reasoning across turns via a new reasoning.context setting.
  • Global rollout on ChatGPT, Codex, and the API, with EU availability confirmed on day one by Ivan Fioravanti in Milan.
  • Live on Cursor within roughly 20 minutes of the livestream, per Ashutosh Shrivastava, and available on Replicate as openai/gpt-5.6-sol the same day.
  • Alongside the model: ChatGPT Work, a redesigned ChatGPT desktop app merging Chat, Work, and Codex, publicly publishable Sites, and a new Plugin Directory replacing the App Directory, per Tibor Blaho's summary.
  • GPT-5.4 sunset date: July 23, 2026. Atlas browser sunset: targeted August 9, 2026.

That is the entire verifiable surface as of publication. Everything past this section is analysis, comparison, or open question.

Why the Naming Change Matters

OpenAI dropped the "Pro / Mini" suffix scheme in favor of celestial tier names. The developer docs are explicit about the reasoning: "The number is the generation; the name is the capability tier, so each tier can advance on its own schedule."

That decoupling is the interesting part. Under the old scheme, gpt-5.5 and gpt-5.5-mini had to move together. With Sol, Terra, and Luna, OpenAI can ship a Sol-only update — say, tighter agentic coordination — without touching Terra or Luna's serving stack. It also means the flagship name survives across generations: the next iteration presumably becomes GPT-5.7 Sol, not GPT-6.

Chinese analyst 歸藏 called out the parallel with Anthropic's approach shortly after the release, noting that OpenAI now names capability tiers the way Claude names model families. Whether that convergence is coincidence or benchmarking against Anthropic's product-line clarity is unstated. Jimmy Apples' wry "needs a cool name" note suggests the naming reshuffle registered in the community as a deliberate strategic move rather than a cosmetic tweak.

The Six Reasoning Levels and Ultra Mode

The reasoning effort axis is now six-wide: none, low, medium, high, xhigh, max. none returns fast, low-latency answers with no visible reasoning traces. max is new, sitting above xhigh, and OpenAI advises comparing both settings on representative workloads before committing — the implication being that max is expensive enough that it should not be a default.

Ultra Mode is separate from reasoning effort. It is a coordination feature, not a depth feature. Where max says "think harder about this one problem," ultra says "split this problem into subagents and run them in parallel, then merge." The release page draws the analogy explicitly to Codex's existing ultra mode.

For builders, this creates a fresh decision axis. A long-horizon refactor benefits from Sol at max. A wide-fan-out research task, where the subtasks are independent, benefits from Sol at ultra. Getting the choice wrong burns tokens without payoff. Simon Willison's pelican benchmark comparison put concrete numbers on the spread: on his SVG-generation test, the cheapest configuration was Luna at effort none for 0.71 cents, the most expensive was Sol at max for 48.55 cents — a 68× cost range across a single family on a single task.

Programmatic Tool Calling and the Multi-Agent Beta

Two API additions have the most far-reaching implications for agent builders.

Programmatic Tool Calling lets GPT-5.6 write JavaScript that orchestrates eligible tool calls, passes results between them, and processes intermediate outputs — all inside a hosted OpenAI runtime, without needing a full turn-by-turn conversation loop back to the model. The developer docs suggest using it for "bounded, tool-heavy workflows that do not require fresh model judgment between each step." That framing matters. It carves out a middle layer between "one tool call per turn" and "a full agent loop," aimed squarely at deterministic pipelines with many tool invocations.

Simon Willison flagged this as reminiscent of Anthropic's dynamic filtering mechanism in Claude's web search tool, where code execution runs over search results within a single model turn.

Multi-agent [beta] is the more ambitious of the two. It lets a single GPT-5.6 instance spin up multiple subagents, run them in parallel, and synthesize their results into a single response. It ships as a beta feature in the Responses API. The obvious use case is research fan-out — asking five subagents to investigate five sub-questions concurrently. The less obvious use case is graders and validators: one subagent generates, one critiques.

Both features push more agentic behavior inside the model call. Historically, agent loops lived in the developer's application code. GPT-5.6 offers to absorb some of that logic into the hosted runtime. Whether builders want that trade — less flexibility for less orchestration overhead — will play out over the next few weeks.

Pricing and the Cost-Efficiency Claim

The pricing table, drawn from Simon Willison's independent verification against the Replicate listing:

TierInput / 1M tokensOutput / 1M tokens
GPT-5.6 Luna$1$6
GPT-5.6 Terra$2.50$15
GPT-5.6 Sol$5$30
Claude Fable 5 (for reference)$10$50
Claude Opus series (for reference)$5$25

Sol matches Opus on input price and beats Fable 5 by roughly 50% on input and 40% on output. Willison correctly notes that per-million-token pricing has become a less useful comparison metric now that reasoning tokens can vary wildly between models for the same task. A cheaper per-token model that burns 3× more reasoning tokens is not cheaper in practice.

OpenAI's own headline efficiency claim, from the release page: "GPT-5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost" on the Artificial Analysis Intelligence Index v4.1. That number reportedly comes from OpenAI's own harness. Independent verification is the natural next step for anyone whose bill scales with this.

For API access, GPT-5.6 is also available on kie.ai as GPT-5.6, callable with a single API key alongside the rest of its model catalog.

GPT-5.6 Sol vs Claude Fable 5: What the Signal Says

The most-cited competitor across the signal bundle is Claude Fable 5, which OpenAI's release page benchmarks against directly. Five comparison dimensions with data from primary sources:

  • Agents' Last Exam: Sol reaches 53.6, described as "13.1 points" above Fable 5 (adaptive reasoning), per the OpenAI release page. That puts Fable 5 at approximately 40.5. This is an OpenAI-published number on an OpenAI-selected benchmark, and should be read as such.
  • SWE-Bench Pro: Willison reports that Claude Fable 5 scored 80% and GPT-5.6 Sol scored 64.6%. OpenAI concurrently published an audit arguing ~30% of SWE-Bench Pro tasks are broken. The timing of that audit release is notable.
  • Artificial Analysis Coding Agent Index: OpenAI claims Sol at max reasoning hits 80, "2.8 points above Fable 5, while using less than half the output tokens." Fable 5 would then sit at ~77.2 on the same index.
  • Input / output pricing per 1M tokens: Sol at $5/$30 vs Fable 5 at $10/$50 — Sol is roughly half the sticker price.
  • Independent third-party benchmarks: unverified — no public number from either lab in this signal set. Community leaderboards will need a few days to catch up.

Willison's own qualitative note is worth quoting: he has had early access to Sol, describes it as "very competent," but adds that "so far it hasn't struck me as better than Fable at the kind of complex coding tasks I've been using with Anthropic's model." That is one experienced tester's impression, not a benchmark, and should be weighted accordingly.

The read of the head-to-head so far: Sol wins on OpenAI's chosen benchmarks and on price. Fable 5 still wins on the one adversarial benchmark (SWE-Bench Pro) that OpenAI is simultaneously trying to invalidate. Both facts can be true.

What We Know vs. What We Don't

What we know:

  • GPT-5.6 ships in three tiers — Sol, Terra, Luna — rolling out globally today across ChatGPT, Codex, and the API, per the OpenAI release page.
  • Pricing per 1M input/output tokens: Luna $1/$6, Terra $2.50/$15, Sol $5/$30, per Simon Willison's notes.
  • Six reasoning effort settings — none, low, medium, high, xhigh, max — per the OpenAI model guidance docs.
  • Ultra Mode coordinates multiple agents across parallel workstreams, per the release page.
  • OpenAI reports Sol reaches 53.6 on Agents' Last Exam, 13.1 points above Fable 5's adaptive-reasoning score on the same benchmark.
  • Per Willison, Sol scored 64.6% on SWE-Bench Pro, Fable 5 scored 80%, and OpenAI simultaneously published an audit questioning ~30% of SWE-Bench Pro tasks.
  • EU availability was confirmed at launch by Ivan Fioravanti — no staged rollout, per his live-from-Milan tweet.
  • The API introduces Programmatic Tool Calling, a Multi-agent beta, explicit prompt cache breakpoints, and persisted reasoning across turns.

What we don't know:

  • Independent third-party benchmark leaderboards for GPT-5.6 are not out — every published number in this bundle is either OpenAI's own or Simon Willison's early-access review.
  • The training data cutoff for GPT-5.6 has not been disclosed in any tweet or link snapshot reviewed here.
  • Context window sizes for Sol, Terra, and Luna are not published in the sources reviewed, though the Replicate endpoint confirms text and image inputs are supported for Sol.
  • Parameter counts remain undisclosed for all three tiers, as is typical for OpenAI releases.
  • Rate limits and burst throughput per tier are not documented in the sources reviewed.
  • The gap between "xhigh" and "max" reasoning is qualitative in the docs — no benchmark separates them cleanly yet.
  • Whether Programmatic Tool Calling can call user-defined tools, or only OpenAI's hosted tools, needs clarification from the developer docs.
  • Ultra Mode's cost curve is not documented — how many parallel subagents it spawns, and how that scales pricing, is unclear.

What Builders Should Do This Week

Three concrete moves before rewiring anything against Sol as a default.

First, rerun your existing eval harness across Terra and Luna, not just Sol. OpenAI is positioning Terra as "GPT-5.5-level quality at roughly half the cost." If that survives contact with your workload, Terra becomes the default and Sol becomes the escalation path. The 68× intra-family cost range Willison observed means tier selection matters more than model selection for most production tasks.

Second, prototype against Programmatic Tool Calling on one bounded pipeline before rewriting agent loops. The natural candidate is any deterministic multi-tool workflow currently living in application code — data validation chains, ETL steps, deterministic scrapers. If it works, you move complexity into the hosted runtime and cut round-trip latency. If it doesn't, you have learned that in an afternoon.

Third, do not re-architect around Multi-agent yet. It ships as beta. Beta at OpenAI has historically meant "the API surface will move." Waiting one release cycle costs little and derisks a lot.

The Week Ahead

Three signals worth tracking:

  • Watch for independent SWE-Bench Verified / SWE-Bench Pro reruns from third parties. OpenAI's audit of SWE-Bench Pro is convenient, and the community will now stress-test whether the audit holds up on the specific tasks where Fable 5 outperformed Sol.
  • Run your own coding eval — not Agents' Last Exam, not SWE-Bench. A representative internal task against Sol vs Fable 5 vs your current baseline. The 13.1-point Agents' Last Exam claim only matters if it survives your workload.
  • Pin the model versions in your pipelines today. GPT-5.4 leaves ChatGPT on July 23; the Atlas browser sunsets around August 9. The auto-migration path is not always what you want.

Building similar long-horizon agentic and coding workflows? On kie.ai you can try GPT-5.6, Claude Opus 5, and Claude Fable 5.

#gpt-5.6#gpt-5.6 sol#gpt-5.6 terra#gpt-5.6 luna#openai release#ultra mode#programmatic tool calling#multi-agent api
Daniel Okonkwo

About Daniel Okonkwo

Daniel writes about inference systems, model architecture, and what new releases actually change for builders.

View all posts by Daniel Okonkwo