What Is Gemini 4 Argon? 1M-Token Output

Daniel Okonkwo

Daniel Okonkwo

Senior ML Engineer

Published: October 6, 2026
Gemini 4 Argon model overview highlighting its 1 million-token output limit

TLDRGemini 4 Argon is Google’s frontier reasoning model with 1M-token output, $2/$10 introductory pricing, and Fairwind-first access.

Inside Gemini 4 Argon: The 1M-Token Output Model

Gemini 4 Argon is Google DeepMind’s frontier reasoning model for long-horizon software engineering, enterprise knowledge work, and defensive cybersecurity. Google announced it on September 30, 2026, and initially limited access to trusted cyber defenders through the Fairwind Program; it is not yet generally available through the Gemini app or public API. Its defining specification is a maximum output of 1 million tokens, up from 64,000 tokens. Introductory pricing is $2 per 1 million input tokens and $10 per 1 million output tokens, according to the official announcement.

Key Takeaways

  • Gemini 4 Argon is a proprietary frontier model built by Google DeepMind.
  • The model can generate up to 1 million output tokens in one response.
  • Initial access is restricted to trusted cyber defenders in Google’s Fairwind Program.
  • Introductory API pricing is $2 per 1 million input tokens and $10 per 1 million output tokens.
  • Google reports 77.9% on DeepSWE v1.1 and 68% on CWE-bench v1.
  • Public API access, rate limits, an official model ID, and the input context window are not yet confirmed.

What Is Gemini 4 Argon?

Gemini 4 Argon is the first publicly announced model in the Gemini 4 generation. Google describes it as a frontier model designed to sustain deep reasoning across complex, multi-step workflows rather than answer only short, isolated prompts.

Its target domains are unusually specific: real-world software engineering, legal and financial knowledge work, research, and defensive cybersecurity. Google is already using Argon internally for coding, infrastructure optimization, technical writing, and quantum-computing research.

Argon is also the public model name, not merely an internal leak label. Google has not explained why it chose “Argon” or how that name maps to familiar Gemini tiers such as Pro, Flash, or Ultra. Claims that a separate Gemini 4 Pro or Flash will follow remain unconfirmed.

The model is officially announced and in a limited rollout. It is neither a general public release nor an undocumented rumor. For the broader generation-level background, see Meet Gemini 4, Google’s Most Ambitious Pre-Training Run Yet.

Gemini 4 Argon at a Glance

A third-party Artificial Analysis model profile lists text and image input with text output. Google has not yet published a complete developer-facing specification sheet.

SpecificationGemini 4 Argon
DeveloperGoogle DeepMind
TypeProprietary frontier reasoning model
ModalityText and image input; text output, per third-party testing
Context windowNot yet confirmed by Google; third-party listing says 1 million tokens
Maximum output1 million tokens
Introductory pricing$2 per 1 million input tokens; $10 per 1 million output tokens
Standard pricing$4 per 1 million input tokens; $20 per 1 million output tokens
Cached input95% discount from input-token price
AvailabilityFairwind Program testers first; paid API and Google AI Ultra planned
LicenseProprietary; detailed license terms not yet published
Model weightsNot available
Public API model IDNot yet confirmed
Announcement dateSeptember 30, 2026

How Gemini 4 Argon Works and What Makes It Different

Google has not published Argon’s parameter count, architecture, training-compute figure, routing design, or full technical report. Its “frontier” designation describes capability and positioning, not a disclosed architecture.

The clearest difference is Million-Token Output. This is an output ceiling, not simply the amount of material the model can read. It gives an agent room to reason, invoke tools, revise work, and produce hundreds of thousands of tokens within one trajectory.

Argon’s defining specification is a 1 million-token output ceiling, not merely a 1 million-token input context.

That ceiling supports what Google frames as a Long-Horizon Workflow: a task that may involve repository exploration, debugging, testing, migration, documentation, and repeated correction. More output capacity does not guarantee correctness, however. It can also raise latency and cost if a task consumes the full allowance.

At introductory rates, one full 1 million-token output would cost $10 before input and caching charges.

The second differentiator is High Reasoning, the configuration used in several third-party evaluations. Artificial Analysis recorded 53 points on its Intelligence Index and a 15% hallucination rate on AA-Omniscience. That result indicates greater willingness to avoid unsupported answers, although the same evaluation placed Argon’s answer accuracy below GPT-6 Astra.

The third differentiator is the Cached Input Discount. Reused prompt content receives a 95% discount from the normal input-token price. This matters for agent systems that repeatedly send the same repository context, policy library, or enterprise knowledge base.

What You Can Do With Gemini 4 Argon

Google’s strongest use cases involve sustained technical work rather than ordinary chat:

  • Repository-scale engineering: Argon scored 77.9% on DeepSWE v1.1, a benchmark for long, multi-step software-engineering tasks.
  • Code migration: Google says Argon agents are migrating C and C++ codebases to Rust, including projects exceeding 800,000 lines of code.
  • Compiler-guided optimization: An Argon workflow replaced 32,000 lines of SIMD code in the libgav1 video decoder. Google says the resulting safe Rust implementation ran 2.7 times faster than the previous Rust port.
  • Infrastructure optimization: Argon agents identified changes that freed more than 300 TiB of memory across Google data centers. Google estimates eventual savings between 500 TiB and 1 PiB.
  • Quantum research: Google reports that Argon improved a published quantum-algorithm baseline by 40% within minutes.
  • Defensive cybersecurity: The model can discover, validate, and help patch vulnerabilities, including generating proof-of-concept evidence for authorized defenders.
  • Enterprise analysis: Legal, finance, tax, research, and business-automation evaluations form a major part of Argon’s benchmark positioning.

These internal deployments are vendor-reported. Google says critical code migrations still undergo automated testing, emulation, manual audits, and human review before production deployment.

Google says Gemini 4 Argon agents have already freed over 300 TiB of memory across its data centers,

Source: @kimmonismus

How Gemini 4 Argon Compares

Argon’s benchmark profile is competitive but not a clean sweep. It leads several coding, long-context, and knowledge-work tests while trailing on some terminal, expert, and agent evaluations. The comparison also mixes vendor-reported results with third-party measurements.

MetricGemini 4 ArgonGPT-6 AstraClaude Opus 5.5
Input price per 1 million tokens$2 introductory; $4 standard$10$4
Output price per 1 million tokens$10 introductory; $20 standard$50$20
Maximum output1 million tokens128,000 tokens128,000 tokens
DeepSWE v1.177.9%74.1%74.2%
Artificial Analysis Intelligence Index53 points, High53 points, Max58 points, Max with fallback
Terminal-Bench 4.0About 57%About 59%About 64% to 66%

Arena placed Gemini 4 Argon High at number 1 in Text Arena with 1,525 points, number 8 in Code Arena: WebDev with 1,679 points, and number 8 in Agent Arena in its initial release update. The Arena announcement also described a blended cost of $8 per 1 million tokens.

Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #

Source: @arena

Those rankings depend on the evaluator and task slice. A later Arena CEO commentary excerpt said Argon fell from number 1 overall to number 7 with expert judges and number 10 on agentic tasks.

Argon is a benchmark leader in several categories, but the evidence does not support calling it the best model for every workload.

For additional context on its main OpenAI comparison point, see the earlier analysis of the GPT-6 Astra release.

Availability: How to Access Gemini 4 Argon

Gemini 4 Argon currently follows a Fairwind-First Rollout. Access began with vetted cyber defenders participating in Google’s Fairwind Program. Google is also participating in the U.S. government’s voluntary pre-release model-access process while testing safeguards.

Paid API customers and Google AI Ultra subscribers are next in the stated rollout sequence. Google has not announced a broader release date, public API endpoint, official model ID, free tier, rate limits, regional coverage, or production service-level terms.

Gemini 4 Argon is not hosted on kie.ai. Teams that need to prototype comparable frontier-chat workflows while Argon remains restricted can evaluate the GPT-6 Astra model page as a separate available option.

Users should avoid unofficial “early access” login pages. Community researchers have already identified impersonation accounts advertising fake access.

What We Don’t Know Yet

Several specifications remain open:

  • Google has not confirmed the official input context window.
  • No public API model ID or endpoint has been published.
  • The duration of the $2/$10 introductory pricing period is unknown.
  • Public rate limits, latency, throughput, and regional availability remain unpublished.
  • Google has not disclosed the parameter count, architecture, training recipe, or knowledge cutoff.
  • The relationship between Argon and possible Gemini 4 Pro or Flash variants is unconfirmed.
  • Independent production testing remains limited because most developers cannot access the model.
  • Google has not published self-hosting rights, weights, or an open-source license.

These gaps matter because a 1 million-token output allowance can create substantial latency and token consumption. Price per token alone does not establish cost per completed task.

Frequently Asked Questions

What is Gemini 4 Argon?

Gemini 4 Argon is Google DeepMind’s proprietary frontier reasoning model for long-horizon software engineering, enterprise knowledge work, and defensive cybersecurity. Its headline specification is a maximum output of 1 million tokens per response.

Is Gemini 4 Argon publicly available?

Gemini 4 Argon is not yet generally available to developers or consumers. Access is limited to trusted cyber defenders through Google’s Fairwind Program, with paid API customers and Google AI Ultra subscribers scheduled to follow.

How much does Gemini 4 Argon cost?

Gemini 4 Argon costs $2 per 1 million input tokens and $10 per 1 million output tokens during its introductory period. Standard pricing is $4 per 1 million input tokens and $20 per 1 million output tokens, while cached input receives a 95% discount.

Does Gemini 4 Argon have a 1M-token context window?

Gemini 4 Argon has a confirmed maximum output of 1 million tokens, but Google has not yet published its input context window. Artificial Analysis lists a 1 million-token context window, so treat that input figure as third-party data until Google’s technical documentation confirms it.

Is Gemini 4 Argon open source?

Gemini 4 Argon is not an open-weights model. Google has announced hosted access through its own programs and future API channels, but it has not released model weights or licensing terms for self-hosting.

Gemini 4 Argon vs GPT-6 Astra: which is better?

Neither model is universally better across every evaluation. Gemini 4 Argon leads the vendor-reported DeepSWE v1.1 result at 77.9% versus GPT-6 Astra’s 74.1%, while both score 53 points on the Artificial Analysis Intelligence Index.

What can Gemini 4 Argon do?

Gemini 4 Argon can handle long-running coding, code migration, research, legal and financial knowledge work, and defensive cybersecurity workflows. Google also uses Argon agents for data-center optimization, quantum algorithm research, and large C/C++-to-Rust migrations.

What to watch next: Google’s public API documentation should settle the model ID, input context window, rate limits, and final modality support. The other critical signals are independent repository-scale evaluations and the date when paid API or Google AI Ultra access begins.

Building similar long-horizon reasoning workflows? On kie.ai you can try Claude Sonnet 5.5, Claude Opus 5.5, and Gemini 3.8 Flash.

Daniel Okonkwo

About Daniel Okonkwo

Daniel writes about inference systems, model architecture, and what new releases actually change for builders.

View all posts by Daniel Okonkwo