Gemini 4 Carbon: Checkpoint Deep Dive & Analysis

Lukas Vogel

Lukas Vogel

Applied Research Editor

Published: October 12, 2026
Illustration of a Gemini 4 internal checkpoint labeled Carbon inside a coding environment

TLDRAn internal Google checkpoint one employee says feels like Opus 5.5 for coding. What the Gemini 4 Carbon signals confirm and what stays open.

Breaking Down Gemini 4 Carbon: What Google's Internal Checkpoint Actually Shows

Gemini 4 "Argon" has not shipped broadly. Yet the model generating the most noise this week is the one nobody outside Google can touch: an internal checkpoint called Carbon, spotted inside Google's own coding platform, that one employee reportedly said "feels like Opus 5.5" for coding.

TLDR Gemini 4 Carbon is an internal Google checkpoint that, per a Business Insider report dated October 9, 2026, staff have been testing on the Jetski coding platform for a few days. The headline claim is a single employee impression that it "feels like Opus 5.5" for coding. There are no published benchmarks, no release date, and no confirmed relationship to the still-limited Argon model. Treat every performance number circulating on X as unverified until Google ships a model card.

Key Takeaways

  • The source of record is a single Business Insider report (byline Hugh Langley), October 9, 2026, citing internal documents and screenshots.
  • Google reportedly tested three Gemini 4 variants internally: Argon, Barium, and Carbon.
  • The public Argon model was internally called Barium-B, so codenames do not map cleanly to public names.
  • The only capability signal is one employee's remark that Carbon "feels like Opus 5.5" for coding — an impression, not a benchmark.
  • Gemini 4 Argon itself was unveiled September 30, 2026, and remained limited to internal teams and the Fairwind program.
  • A Google DeepMind employee publicly nodded to recursive self-improvement as a reason for the fast iteration.

What Was Actually Reported

Strip away the amplification and the primary source is narrow. According to Business Insider's October 9 report, which cites documents, screenshots, and internal chats, Google staff have been using "a new version of Gemini 4" dubbed Carbon on Jetski, the company's internal coding platform, "over the last few days."

Three internal comments circulated, all pulled from messaging channels rather than any official statement: one that Carbon "feels comparable to Opus 5.5," another that "Carbon is really good!", and a third describing a "new Gemini pro next model." The employee behind the most-quoted line added a caveat that more testing was needed before a verdict.

That is the whole of the first-party evidence. Everything else is relay. Of the nine public posts we reviewed, most repackage the same Business Insider-attributed claim. Chris laid out the rollout framing, noting that staff are "testing better versions that may further close the gap with rivals."

Google is preparing to roll out its new “Gemini 4 'Argon' model to the public”, It goes onto say sta

Source: @ChrisGPT

Harshith summarized the naming ladder: internal names Argon, Barium, and Carbon, with public Argon reportedly being the Barium-B checkpoint rather than the newest one.

The Argon–Barium–Carbon Naming Ladder

The naming is where most of the confusion starts, so it is worth pinning down. The Naming Ladder is simple to state and easy to misread: Argon, Barium, and Carbon are internal Gemini 4 codenames, and the public "Argon" model was internally labeled Barium-B.

That single detail breaks the tempting assumption that the codenames form a public product sequence. An internal Barium checkpoint became public-facing Argon. By the same logic, Carbon could ship under the Argon name as an update, or it could arrive as a distinct model. Business Insider could not determine which, and Google declined to comment.

A useful reframing circulated on Threads. According to a post from testingcatalog, Jetski — the internal coding tool named in the report — is itself an internal name for Antigravity, and Google has added low, medium, and high reasoning-effort references to the Argon model there. If accurate, that places Carbon inside Google's agentic coding tooling rather than a consumer surface.

So the honest headline is narrower than "Google is launching Gemini 4 Carbon." A development candidate surfaced in an internal tool. That is the event.

The "Feels Like Opus 5.5" Claim, Examined

The line doing all the work is a vibe check. "Feels like Opus 5.5" is an employee's subjective impression of coding quality, offered without a benchmark, a task set, or an evaluator description.

Context matters here. Per the same reporting, earlier Argon versions reminded another employee of the older Opus 5 on some coding tasks. If the internal trajectory ran from an Opus 5-ish Argon to an Opus 5.5-ish Carbon, that is a plausible iteration story, but it is still an anecdote against a moving target. Anthropic's Opus 5.5 is the model developers currently reach for on long-horizon agentic coding, where the system works across many files and steps on its own. Comparing to it is a compliment, not a measurement.

A lot is happening at Google right now. Gemini 4 "Argon" hasn't even been released yet, and internal

Source: @kimmonismus

One claim went further than the source supports. Salio posted that Carbon is "Beating Opus 5.5 and GPT-6 Astra internally" and tied it to recursive self-improvement. No published number backs that. On the limited evidence so far, Gemini 4 Carbon appears to impress Google's own engineers on coding — but whether that survives contact with an independent SWE-bench or DeepSWE run is entirely unverified.

The Recursive Self-Improvement Angle

The most interesting thread is not the benchmark bragging; it is the cadence. Google unveiled Argon on September 30, and roughly nine days later staff were already testing a reportedly stronger checkpoint. That pace is what prompted a Google DeepMind employee, Vedant Misra, to respond to the report on X by referencing recursive self-improvement — using AI to build better AI.

Recursive Self-Improvement (RSI) is the hypothesis that a lab can compress model-iteration cycles by using its current models to help design, train, or evaluate the next ones. Google has done something similar with its Flash cadence, and both OpenAI and Anthropic have reported comparable internal gains. The claim is not new. What is notable is a named insider invoking it publicly to explain a specific checkpoint cadence.

Treat RSI as framing, not proof. A fast internal checkpoint is consistent with RSI and also consistent with ordinary fine-tuning iterations. The signal does not let you distinguish the two.

Gemini 4 Carbon vs Claude Opus 5.5: What the Reports Say

The comparison is unavoidable because the source made it, so here is what each dimension actually rests on.

  • Coding impression. Carbon: one employee said it "feels like Opus 5.5," an anecdotal internal remark. Opus 5.5: the reference model developers use today for long-horizon agentic coding. This dimension is a vibe check, not a head-to-head.
  • Published benchmark. Carbon: unverified — no public number in available sources. Opus 5.5: a third-party writeup citing Google's own Argon launch table lists 74.2% on DeepSWE v1.1 for Opus 5.5, against 77.9% for Argon and 74.1% for GPT-6 Astra. Those are vendor figures for Argon, not Carbon, and not reproduced independently.
  • Availability. Carbon: internal-only, tested on Jetski. Opus 5.5: publicly available and, per the reporting, already brought into Google's Antigravity for some subscriber tiers.
  • Context window. Both: unverified — no public number from either lab so far.

The one measured-looking number in the whole cluster, 77.9% DeepSWE v1.1, belongs to Argon and comes from Google's own launch table. It says nothing directly about Carbon. At best it anchors the family Carbon reportedly improves on. Anyone quoting it as a Carbon score is misreading the chain.

What We Know vs. What We Don't

What we know:

  • Business Insider reported on October 9, 2026 that Google staff have been testing an internal Gemini 4 checkpoint called Carbon on the Jetski coding platform over the last few days.
  • One Google employee reportedly said Carbon "feels like Opus 5.5" for coding, an early internal impression rather than a published benchmark, with the caveat that more testing was needed.
  • An internal document reviewed by Business Insider lists three Gemini 4 names: Argon, Barium, and Carbon.
  • The model selected for the public Argon release was internally called Barium-B, so public codenames do not map one-to-one to internal checkpoints.
  • Gemini 4 Argon was unveiled on September 30, 2026 and remained limited to Google's internal teams and the Fairwind program rather than a broad public release.
  • A Google DeepMind employee, Vedant Misra, responded to the report on X by referencing recursive self-improvement as a possible driver of the rapid iteration.

What we don't:

  • Whether Carbon will ship as a standalone Gemini 4 model or as another update inside the Argon family; Business Insider could not determine this and Google declined to comment.
  • Any official benchmark for Carbon; none have been published, and the only capability signal is anecdotal.
  • Carbon's release date, final name, pricing, and context window, all of which are unconfirmed.
  • Whether "feels like Opus 5.5" reflects measured performance; no benchmark, task results, or evaluator details were provided.
  • Whether the single X claim that Carbon beats Opus 5.5 and GPT-6 Astra internally is accurate; it is unverified and single-source.

Why This Matters for Builders

The practical takeaway is about discipline, not Google. A positive internal impression on a coding checkpoint is the weakest tier of evidence a builder can act on. It is unlabeled, unreproduced, and selected by the people who made the model.

If you ship anything that depends on a frontier coding model, the Carbon story is a reminder to separate roadmap awareness from procurement decisions. Knowing Google is iterating fast is useful for planning. Swapping a production coder on the strength of a "feels like" quote is not. The gap between those two is exactly the gap between an internal checkpoint and a model card.

There is also a naming lesson. If Barium-B became public Argon, then any future "Carbon" you see in an API string may not be the Carbon from this report. Pin behavior to evals, not to codenames.

How to Evaluate It Yourself When It Ships

When a Gemini 4 coding model does reach a public endpoint — whether it carries the Argon name or something new — the Carbon reports give you a ready-made test plan.

First, reproduce the comparison that started this. Run the same long-horizon agentic coding tasks against both the Google model and Opus 5.5, on your own repositories, and record pass rates rather than impressions. Second, check the reasoning-effort controls. If the Antigravity references to low, medium, and high effort hold, cost and latency will swing hard across those settings, so benchmark the setting you will actually run in production. You can prototype that harness today against a comparable public model such as Gemini 3.1 Pro and simply repoint it when the Gemini 4 coder lands.

Third, ignore the 77.9% until Google publishes a Carbon-specific number. A single vendor DeepSWE figure for Argon is a family hint, not a Carbon result.

What To Watch Next

Three concrete signals will tell you whether Carbon is real product movement or internal noise. Watch for an official Gemini 4 model card or API changelog that names Carbon or confirms an Argon update with new benchmarks. Run your own coding eval before trusting any "beats Opus 5.5" framing, since every performance claim in this cluster is single-source. And track whether the public Argon rollout finally widens beyond the Fairwind program, because a stalled Argon launch alongside a hyped internal Carbon would say more about Google's release strategy than about either model.

Building similar coding and agentic-reasoning workflows? On kie.ai you can try Claude Opus 5.5, GPT-6 Astra, and Kimi K3.

#gemini 4 carbon#gemini 4 argon#google deepmind checkpoint#opus 5.5 coding#jetski coding platform#ai model leak analysis
Lukas Vogel

About Lukas Vogel

Lukas reads the papers and model cards so you do not have to, focusing on reproducible claims.

View all posts by Lukas Vogel