What Is Gemini 3.6 Flash? Pricing, Benchmarks & Availability

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: July 21, 2026
Gemini 3.6 Flash model identifier surfaced inside the Google Antigravity IDE

TLDRGemini 3.6 Flash is Google's Flash-tier model, launched July 21, 2026 at $1.50 / $7.50 per 1M tokens with a 1M context window and day-one availability in AI Studio, the Gemini API, and the Gemini app.

Gemini 3.6 Flash 101: The gemini-3.6-flash-tiered Model ID Spotted Inside Google Antigravity

Gemini 3.6 Flash is Google's newest Flash-tier chat model, launched on July 21, 2026 after its gemini-3.6-flash-tiered identifier first surfaced inside the Google Antigravity IDE earlier the same day. Google confirmed the model with an official blog post from the Gemini team, a model card, and day-one availability in Google AI Studio, the Gemini API, the Gemini app, Android Studio, Antigravity, and Vertex AI / Gemini Enterprise. Pricing is $1.50 per 1M input tokens and $7.50 per 1M output tokens — a lower output rate than Gemini 3.5 Flash — with a 1M-token context window and roughly 17% fewer output tokens on the Artificial Analysis Index than its predecessor.

Key Takeaways

  • Gemini 3.6 Flash launched on July 21, 2026 alongside Gemini 3.5 Flash-Lite and a limited-pilot Gemini 3.5 Flash Cyber model.
  • The public model ID is gemini-3.6-flash; the gemini-3.6-flash-tiered string first spotted in Google Antigravity by user @ChrisGPT (source) was an internal preview label.
  • Pricing is $1.50 per 1M input tokens and $7.50 per 1M output tokens, versus $1.50 / $9.00 for Gemini 3.5 Flash. Cached input is $0.15 per 1M tokens.
  • The context window is 1,048,576 input tokens and up to 65,536 output tokens, with multimodal input (text, image, video, audio, PDF) and text output.
  • On Google's own benchmarks, 3.6 Flash beats 3.5 Flash on DeepSWE (49% vs 37%), OSWorld-Verified (83.0% vs 78.4%), MLE-Bench (63.9% vs 49.7%), and GDPval-AA v2 (1421 vs 1349 Elo), while using about 17% fewer output tokens on the Artificial Analysis Index.
  • Independent benchmarks are more mixed: the Artificial Analysis Intelligence Index gives it a score of 50, matching Gemini 3.5 Flash, and community tests show it trailing GPT-5.6 Luna, Grok 4.5, and Kimi K3 on several coding evaluations.
  • Gemini 3.5 Pro is still in partner testing, and Google confirmed it has started pre-training Gemini 4.

What Is Gemini 3.6 Flash?

Gemini 3.6 Flash is a Flash-tier large language model from Google DeepMind, officially launched on July 21, 2026. Its identifier was first evidenced by community users inside Google Antigravity earlier that day, before Google's own announcement went live. Antigravity, Google's agentic IDE, exposed gemini-3.6-flash-tiered in its selector, and screenshots spread through AI-watcher accounts on X within hours.

"Gemini 3.6 flash spotted in antigravity! However they have been testing this model for a few days!" wrote user @ChrisGPT in the post that concentrated most of the early discussion.

Gemini 3.6 flash spotted in antigravity! However they have been testing this model for a few days! B

Source: @ChrisGPT

Later the same day, Google AI Studio announced the model publicly, describing it as "our latest model that balances speed with intelligence to deliver strong performance in agentic and multimodal tasks," built directly on developer and customer feedback from Gemini 3.5 Flash. The Gemini team's blog post positioned 3.6 Flash as the new "workhorse" of the Flash line — better on coding, knowledge work, and multimodal tasks, with meaningfully lower token usage per task and a reduced output-token price.

The label follows Google's established naming convention for the Gemini family. Numbers such as 3.5 and 3.6 track the generation; tier words like Flash and Pro track capability. Gemini 3.6 Flash sits as an incremental follow-on to Gemini 3.5 Flash rather than a next-generation frontier release — a pattern Google underscored by noting that Gemini 3.5 Pro remains in partner testing and that pre-training has begun on Gemini 4.

Gemini 3.6 Flash at a Glance

AttributeValue
DeveloperGoogle DeepMind
TypeChat / LLM (Flash tier)
Model identifiergemini-3.6-flash (public); gemini-3.6-flash-tiered seen in Antigravity pre-launch
ModalityInput: text, image, video, audio, PDF. Output: text
Context window1,048,576 input tokens / 65,536 output tokens
Pricing$1.50 / 1M input tokens, $7.50 / 1M output tokens, $0.15 / 1M cached input
AvailabilityGoogle AI Studio, Gemini API, Gemini app, Android Studio, Antigravity, Vertex AI / Gemini Enterprise
AnnouncementGoogle DeepMind / Gemini team blog, July 21, 2026
PredecessorGemini 3.5 Flash (1M context, $1.50 / $9 per 1M tokens)
LicenseProprietary

How Gemini 3.6 Flash Works and What Makes It Different

Google has not disclosed architecture details, but it has been specific about the improvements over Gemini 3.5 Flash.

The headline change is efficiency. According to Google, Gemini 3.6 Flash consumes about 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, with reductions of up to 65% on some DeepSWE by Datacurve tasks. It also takes fewer reasoning steps and tool calls to complete multi-step workflows, which — combined with the lower per-token output price — reduces the total cost per agentic task. Jeff Dean highlighted the token-efficiency gain in a side-by-side demonstration on launch day.

The -tiered suffix that appeared on the model id in Antigravity before launch turned out to be an internal preview label. The public API model ID is gemini-3.6-flash, documented on ai.google.dev and in the Gemini Enterprise Agent Platform. The venue where the id first appeared is still notable: Antigravity is Google's agentic coding IDE, and its model selector has become a semi-regular preview channel — earlier in 2026 it exposed anonymous Flash checkpoints on LM Arena that community trackers linked to unreleased Gemini variants. Community user @Lentils80 reported the id appeared in Antigravity "minutes earlier" alongside first tests, and Patrick Loeber of the Gemini team confirmed the team had been building with 3.6 Flash internally "over the last couple of weeks."

On the capability side, computer use is now a built-in client-side tool available via the Gemini API and Gemini Enterprise — a change from 3.5 Flash, where the feature was added mid-cycle. The model also supports thinking, function calling, structured outputs, code execution, URL context, search and Maps grounding, file search, and context caching, and it can be consumed via the Batch API, Flex inference, or Priority inference.

What You Can Do With Gemini 3.6 Flash

Google positions Gemini 3.6 Flash as its workhorse Flash-tier model for coding, knowledge work, and agentic workflows, with three concrete use-case categories highlighted at launch:

  • Coding and code migrations. In Google's own benchmarks, 3.6 Flash jumps to 49% on DeepSWE (from 37% for 3.5 Flash) and shows lower compile-failure and revision rates across building, prototyping, and IDE agent environments. Google demoed it executing code migrations via multi-agent orchestration with lower latency and higher quality than 3.5 Flash.
  • Knowledge work and document analysis. GDPval-AA v2 rises to 1421 Elo from 1349, and Google says customers like Hebbia and Harvey have found the model particularly capable at document parsing, chart and data analysis, and report drafting. MLE-Bench improves to 63.9% from 49.7% for ML research tasks.
  • Computer use and agentic UI control. OSWorld-Verified climbs to 83.0% from 78.4%, and computer use is now built in as a client-side tool via the API.

Real-world reception is more mixed than Google's numbers alone suggest. On the Artificial Analysis Intelligence Index, Gemini 3.6 Flash scores 50 — matching Gemini 3.5 Flash — and Bindu Reddy's benchmark had the new version scoring below its predecessor, which she called "the first time we have seen this in the history of our benchmark." Mark Kretschmann's read of the released numbers was that it loses most coding benchmarks and GDPVal to Grok 4.5 and GPT-5.6 Luna while costing more per output token, with clear strengths on MLE-Bench, OSWorld, chart reasoning, and long context. For a Google flagship comparison, see our profile of Gemini 3.5 Pro, which is still in partner testing at launch.

How Gemini 3.6 Flash Compares

The clearest anchor is the previous Flash generation.

ModelContextPricing (input / output per 1M)Status
Gemini 3.6 Flash1M tokens (65k output)$1.50 / $7.50Launched July 21, 2026
Gemini 3.5 Flash1M tokens$1.50 / $9.00Superseded in the Gemini app by 3.6 Flash
Gemini 3.5 Pro2M tokens (reported)Not yet publishedIn partner testing, unreleased

Against competitors, community coverage placed Gemini 3.6 Flash in the same benchmark cycle as Grok 4.5, GPT-5.6 Luna, and Kimi K3. Chubby (@kimmonismus) noted it is "priced between" GPT-5.6 Luna and Sonnet 5 and "generally performs better," calling it "a solid mid-tier release." On DeepSWE, LinearUncle found it "roughly comparable to Opus 4.8 Medium," ahead of GLM-5.2 and close to Grok 4.5. For the open-weights comparison at that tier, see our writeup on Kimi K3, which ships with 1M context and a documented $3 / $15 per 1M price.

Availability: How to Access Gemini 3.6 Flash

Gemini 3.6 Flash is generally available as of July 21, 2026. Google confirmed the following access paths on launch day:

  • Google AI Studio and the Gemini API — as gemini-3.6-flash. Documentation is live at ai.google.dev, and the model supports the standard Gemini feature set including thinking, structured outputs, function calling, computer use, and caching.
  • The Gemini app — Gemini 3.6 Flash has replaced Gemini 3.5 Flash in the model selector, per Mark Kretschmann and Google's own announcement.
  • Google Antigravity — where the identifier was first spotted, and which now exposes the launched model to users of the IDE.
  • Android Studio — as a selectable model for AI-assisted coding.
  • Vertex AI and Gemini Enterprise — under the Gemini Enterprise Agent Platform, with support for Provisioned Throughput, Batch inference, and Pay-as-you-go tiers.

Third-party integrations followed almost immediately, including OpenCode and GitHub Copilot's model catalog.

If you want to build Flash-tier agentic workloads on models available through kie.ai's catalog, options include the live Gemini 3.5 Flash and other production chat models with documented pricing and quotas.

What We Don't Know Yet

Most of the launch-day unknowns have now been resolved by Google's model card and blog post. A few open items are still worth tracking:

  • True knowledge cutoff. Google's model documentation lists a July 2026 latest-update date, and community reports on launch day cited a March 2026 knowledge cutoff, but some early hands-on tests suggested the model's practical knowledge lagged that claim. Independent verification is still emerging.
  • Independent long-run benchmarks. Beyond Artificial Analysis and Google's own numbers, third-party evaluations of agentic reliability, tool-use stability, and real-world cost per task are still landing. Early community results are notably more skeptical than Google's marketing.
  • Relationship to Gemini 3.5 Pro. Google says 3.5 Pro is still testing with partners and will ship "as soon as it's ready," but the public timeline remains unstated. Community speculation, including from @mark_k on July 18, 2026 (source), floated the theory that Google could reshuffle the naming across the delayed Pro and the emerging Flash. That theory did not play out — 3.6 Flash launched under its expected name — but the Pro release date is still open.
  • Gemini 4 timeline. Google confirmed it has started its "most ambitious pre-training run yet" for Gemini 4. Availability, capability scope, and rollout are all unknown.

Frequently Asked Questions

Is Gemini 3.6 Flash released?

Yes. Google officially launched Gemini 3.6 Flash on July 21, 2026 alongside Gemini 3.5 Flash-Lite and a limited-pilot Gemini 3.5 Flash Cyber model. Gemini 3.6 Flash is generally available in Google AI Studio, the Gemini API, the Gemini app, Android Studio, Antigravity, and Vertex AI / Gemini Enterprise.

Where was Gemini 3.6 Flash spotted?

Before its official launch, Gemini 3.6 Flash was spotted inside Google Antigravity, Google's agentic IDE, on July 21, 2026. Community accounts including @ai_for_success and @ChrisGPT posted screenshots showing the model selector exposing the id gemini-3.6-flash-tiered. Google formally announced the model later the same day and rolled it out across AI Studio, the Gemini API, and the Gemini app.

How much does Gemini 3.6 Flash cost?

Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, with a cached-input rate of $0.15 per million tokens. That is a lower output price than Gemini 3.5 Flash's $9.00 per million output tokens; the input price is unchanged.

What does the tiered suffix in gemini-3.6-flash-tiered mean?

The public model ID at launch is gemini-3.6-flash. The gemini-3.6-flash-tiered string seen in Antigravity before launch was an internal preview label; Google has not documented it as a separate product. Use gemini-3.6-flash in the Gemini API and Google AI Studio.

Is Gemini 3.6 Flash better than Gemini 3.5 Flash?

On Google's official benchmarks Gemini 3.6 Flash improves on Gemini 3.5 Flash across coding, agentic, and multimodal tests — for example DeepSWE 49% vs 37%, OSWorld-Verified 83.0% vs 78.4%, MLE-Bench 63.9% vs 49.7%, and GDPval-AA v2 1421 vs 1349 Elo — while using about 17% fewer output tokens on the Artificial Analysis Index and up to 65% fewer on some DeepSWE tasks. Independent benchmarks are more mixed: the Artificial Analysis Intelligence Index gives it a score of 50, matching 3.5 Flash, and community tests show it trailing peers like GPT-5.6 Luna, Grok 4.5, and Kimi K3 on several coding evaluations.

When will Gemini 3.6 Flash be released?

Gemini 3.6 Flash was released on July 21, 2026. It began rolling out in the Gemini API the same day and took over from Gemini 3.5 Flash in the Gemini app. Gemini 3.5 Flash-Lite launched on the same date, and Gemini 3.5 Flash Cyber is being released as a limited pilot inside Google DeepMind's CodeMender agent for trusted partners and governments.

Is Gemini 3.6 Flash open source?

No. Gemini 3.6 Flash is proprietary and served via Google's APIs and products. Google has never open-sourced a Gemini frontier model.

What to Watch Next

Three concrete signals to track now that the launch itself is settled. First, how independent long-run benchmarks and real-world agent deployments compare to Google's officially reported numbers — the early split between Google's benchmarks and the Artificial Analysis Intelligence Index (which shows no gain over 3.5 Flash) is the most important open question. Second, when Gemini 3.5 Pro exits partner testing and what its capability and pricing look like relative to competitors that have already shipped. Third, any concrete signals on the Gemini 4 pre-training run Google confirmed at launch. This page will be amended as each of those lands.

Building similar Flash-tier chat and agent workloads? On kie.ai you can try Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair