Gemini 3.8 Flash Is a Cost-Focused Workhorse — Its 1M-Token Context Claim

Maya Chen

Maya Chen

Lead AI Researcher

Published: August 30, 2026
Gemini 3.8 Flash reference page covering internal testing, context claims, pricing, and availability

TLDRGemini 3.8 Flash officially launched on September 2, 2026, with a 1-million-token context window, $0.75/$3.75 introductory pricing, and access through the Gemini API, AI Studio, Gemini app, and developer platforms.

Gemini 3.8 Flash is Google's generally available Flash model, officially launched on September 2, 2026. Google positions it as a workhorse for long-horizon software engineering, autonomous agents, and complex enterprise workflows, with improvements over Gemini 3.7 Flash in coding, agentic tasks, and multi-step reasoning. The pre-launch internal preview was reported by LuminaBench on August 28, 2026, but the model is now publicly documented and accessible through the Gemini API and Google AI Studio.

The official API model page lists the stable model ID gemini-3.8-flash, a 1,048,576-token input limit, a 65,536-token output limit, and support for text, images, video, audio, and PDF inputs. Gemini 3.8 Flash launched at $0.75 per 1 million input tokens and $3.75 per 1 million output tokens at its introductory price.

Key Takeaways

  • Gemini 3.8 Flash launched on September 2, 2026, and is generally available.
  • Google describes it as its most intelligent Flash workhorse model, focused on coding, autonomous agents, and specialized multi-step reasoning.
  • The stable model ID is gemini-3.8-flash, with a 1,048,576-token input limit and a 65,536-token output limit.
  • Inputs include text, images, video, audio, and PDF files; the model produces text output.
  • Introductory pricing is $0.75 per 1 million input tokens and $3.75 per 1 million output tokens through the end of 2026.
  • Gemini 3.8 Flash is available through the Gemini API and Google AI Studio, and it is also present in the Gemini app for Google AI Pro and Ultra users.
  • Supplied benchmark results include 73.7% on DeepSWE 1.1, 54.9% on HLE-Verified, and a score of 59 on the Artificial Analysis Intelligence Index.
  • The model can use more output tokens on difficult tasks because it takes additional reasoning steps and verifies its work more often.

What Is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google's latest generally available Flash model. It belongs to the fast-moving Gemini Flash family, which Google describes as an agentic workhorse tier between deeper-reasoning Pro models and lower-cost, high-throughput Flash-Lite models.

Google officially introduced Gemini 3.8 Flash on September 2, 2026, alongside Gemini 3.8 Flash Cyber, a separate variant focused on cybersecurity. The announcement describes 3.8 Flash as Google's most intelligent workhorse model, with significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical multi-step reasoning in specialized domains. Google's launch announcement is available here.

The model is no longer an internal preview or a speculative future release. It is generally available through the Gemini API and Google AI Studio. Developers can use the stable model ID gemini-3.8-flash, while Google Cloud documentation also lists it as a GA model for enterprise agent workflows.

Gemini 3.8 Flash is a launched speed-and-cost workhorse for coding, agents, and long-running tasks—not an unreleased model candidate.

The launch followed a rapid Flash cadence. Community timelines place Gemini 3.6 Flash on July 21, 2026, and Gemini 3.7 Flash on August 13, 2026. Google described Gemini 3.8 Flash as its third Flash release in six weeks. The release-cadence post is available on X.

For background, the earlier Gemini 3.7 Flash overview explains the Flash family’s workhorse positioning and the short interval between recent versions.

Gemini 3.8 Flash at a Glance

SpecificationCurrent information
DeveloperGoogle
Launch statusGenerally available
Launch dateSeptember 2, 2026
Stable model IDgemini-3.8-flash
TypeFlash-series workhorse model
Input modalitiesText, image, video, audio, and PDF
Output modalityText
Context window1,048,576 input tokens, approximately 1 million tokens
Maximum output65,536 tokens
Thinking levelsLow, medium, and high; minimal is not supported
Pricing$0.75 per 1 million input tokens and $3.75 per 1 million output tokens at the introductory rate
Gemini APIAvailable
Google AI StudioAvailable
Gemini appAvailable for Google AI Pro and Ultra users
Google SearchAvailable in AI Mode for Google AI Pro and Ultra subscribers
Developer platformsAvailable in Google Cloud Agent Studio, Vertex AI, Antigravity, OpenRouter, and Cursor
Open weightsNot identified in the supplied material
LicenseNot identified in the supplied material

The 1-million-token context claim that appeared in pre-launch community reporting is now supported by Google's official API documentation, which lists an input token limit of 1,048,576. The same page lists a maximum output limit of 65,536 tokens. The official Gemini API model page provides the specification.

The model accepts text, images, video, audio, and PDF files as inputs. It supports caching, code execution, computer use in preview, file search, function calling, Google Maps grounding, search grounding, structured outputs, URL context, and thinking at low, medium, and high levels.

How Gemini 3.8 Flash Works / What Makes It Different

Google has not published a complete architecture description for Gemini 3.8 Flash. The supplied material does not establish its parameter count, tokenizer design, training method, or internal model architecture. Those details remain open.

Google has, however, described the model's operating behavior. On complex tasks, Gemini 3.8 Flash takes additional reasoning steps and calls tools iteratively. It also verifies its work more often. This design can improve the quality of long-running coding and agentic tasks, but it can increase token usage.

The official Cloud guide describes a trade-off with Gemini 3.7 Flash: 3.8 Flash provides better accuracy and more reliable performance, while consuming more tokens. Developers can use thinking levels to control that trade-off. Lower effort levels can reduce compute overhead for latency-sensitive tasks, while higher levels allow the model to spend more effort on difficult work.

Five practical labels describe the released model:

  • Long-Horizon Coding: Google built the model for software-engineering tasks that extend across multiple steps and tool calls.
  • Agentic Workhorse: The model is intended to plan, act, inspect results, and continue working toward a goal.
  • Iterative Verification: The model can take smaller steps and check its work more frequently.
  • Multimodal Processing: The API accepts text, image, video, audio, and PDF inputs.
  • Effort Control: Low, medium, and high thinking levels let developers adjust the balance between reasoning effort, token use, and latency.

These labels describe documented capabilities and positioning rather than separate product names. The model's architecture and parameter count remain undisclosed, but its public API behavior and supported tools are now documented.

The important post-launch distinction is not whether Gemini 3.8 Flash exists, but how its extra reasoning and verification affect cost, quality, and task completion in production.

What You Can Do With Gemini 3.8 Flash

Gemini 3.8 Flash is available for direct testing through Google AI Studio and for application development through the Gemini API. It is also available in Google Cloud Agent Studio, Vertex AI, Antigravity, OpenRouter, and Cursor.

Coding and code generation

Software engineering is the clearest announced use case. Google describes Gemini 3.8 Flash as engineered for long-horizon software engineering, and the launch material reports substantial gains over Gemini 3.7 Flash on DeepSWE 1.1.

The model can support code generation, debugging, repository work, test creation, command-line tasks, and iterative changes. A supplied Cursor benchmark reported 69.9% performance at a reported cost of $2.38 per task. Those results indicate that the model is aimed at practical coding agents rather than only short code-completion prompts.

Agentic task loops

Gemini 3.8 Flash is designed for multi-step agent workflows. It can use function calling, code execution, file search, URL context, Google Search, Google Maps, and computer use in preview.

Managed Agents provide a direct way to test the model in an agentic environment. A single API call can provision a remote sandbox with Python, Node.js, Git, Bash, network access, filesystems, background support, and selected tools. The supplied Managed Agents guidance describes this environment as a way to run and continue long-running agent interactions.

Browser and terminal assistance

The model's supported tools make it suitable for research and command-line workflows. Search grounding, URL context, computer use, code execution, and function calling can be combined with the model's reasoning process.

Terminal-Bench 2.1 measures whether an agent can complete difficult command-line and coding tasks end to end. Google's Cloud developer guide lists a 90.8% result for Gemini 3.8 Flash on that evaluation, while supplied launch commentary also cites a 73.7% result on DeepSWE 1.1 for long-horizon software engineering.

Tool reliability still needs to be assessed in individual deployments. Benchmark scores do not establish how the model will behave with every permission structure, repository, prompt-injection risk, or recovery scenario.

High-volume interactive applications

The Flash tier is intended for workloads where response time and token cost accumulate quickly. Gemini 3.8 Flash combines the relatively low introductory price of $0.75 per 1 million input tokens and $3.75 per 1 million output tokens with high measured output speed.

Artificial Analysis lists the high reasoning setting at roughly 305 output tokens per second and gives the model a score of 59 on its Intelligence Index. Developers can also use lower thinking levels for tasks where lower compute use or faster responses matters more than maximum effort.

How Gemini 3.8 Flash Compares

The launch evidence now contains direct comparisons with Gemini 3.7 Flash and Claude Opus 5. The comparisons show strong performance in several coding, reasoning, finance, legal, and agent evaluations, but they do not establish that Gemini 3.8 Flash wins every workload.

ModelEvidence statusReported positioningPublished price in supplied evidence
Gemini 3.8 FlashGenerally available since September 2, 2026Long-horizon coding, autonomous agents, multimodal processing, and specialized reasoning$0.75 per 1 million input tokens and $3.75 per 1 million output tokens at the introductory price
Gemini 3.7 FlashGenerally available before the 3.8 launchEarlier Flash workhorse for coding and agents$0.75 per 1 million input tokens and $3.75 per 1 million output tokens at its introductory price
Claude Opus 5Competitor model discussed in supplied comparisonsHigher-cost frontier reasoning and coding$5 per 1 million input tokens and $25 per 1 million output tokens

On the Artificial Analysis Intelligence Index, Gemini 3.8 Flash scored 59. The supplied research summary describes this as a three-point improvement over Gemini 3.7 Flash. Artificial Analysis also lists separate low, medium, and high versions, with the high setting at 59, medium at 57, and low at 52.

On HLE-Verified, a supplied comparison lists Gemini 3.8 Flash at 54.9% versus 54.4% for Claude Opus 5. On Harvey's Legal Agent Benchmark, another supplied comparison lists Gemini 3.8 Flash at 10.0% versus 6.7% for Opus 5. These are specific evaluation results, not a general guarantee that Gemini will outperform Opus in every legal or reasoning task.

Agent Arena placed Gemini 3.8 Flash High at #14 overall with a +5.94% net improvement, compared with Gemini 3.7 Flash High at #32 and +0.84%. Arena's separate cost report lists a median cost of $0.22 per task for Gemini 3.8 Flash High.

The reported advantage comes with a token-use trade-off. One supplied comparison says Gemini 3.8 Flash used about 1.34 times as many output tokens as Gemini 3.7 Flash on DeepSWE, with 143,000 versus 107,000 output tokens. That extra work can improve results, but it matters when evaluating total application cost rather than only per-token pricing.

The earlier Gemini 3.7 Flash pricing analysis provides additional background on the preceding Flash model's pricing.

Availability: How to Access Gemini 3.8 Flash

Gemini 3.8 Flash launched on September 2, 2026, and is generally available.

Developers can use the stable model ID gemini-3.8-flash through the Gemini API. Google AI Studio also provides access, and the supplied Managed Agents announcement identifies a free tier for trying Gemini 3.8 Flash in an agentic environment.

For cloud developers, the model is listed in Google Cloud Agent Studio and Vertex AI. Philipp Schmid described it as the default model in Antigravity and Managed Agents, while OpenRouter announced that Gemini 3.8 Flash was live on its platform. Cursor separately announced availability in its coding environment.

Consumer access is also available. Google Search has Gemini 3.8 Flash in AI Mode for Google AI Pro and Ultra subscribers around the world, and the model is available in the Gemini app for Pro and Ultra users.

The official API documentation lists support for caching, code execution, computer use in preview, file search, function calling, Google Maps grounding, search grounding, structured outputs, thinking, and URL context. Live API, audio generation, and image generation are not supported on the documented model page.

For related coverage of the earlier Flash release, developers can review Gemini 3.7 Flash on kie.ai.

What We Don't Know Yet

The launch resolved the main pre-release questions, but several technical and operational questions remain open:

  • Architecture and parameter count: Google has not published a confirmed parameter count, tokenizer description, or full architecture explanation.
  • Independent generalizability: The supplied benchmarks are promising, but early independent testing is still limited. One repository test found Gemini 3.8 Flash High fixed 20 of 105 hidden bugs at a reported cost of $9.78, while other comparisons produced different rankings and costs.
  • Benchmark methodology: The supplied results come from different evaluations, settings, harnesses, and test conditions. They should not be combined into one universal ranking.
  • Token efficiency: The model can use more output tokens than Gemini 3.7 Flash and other compared models on some tasks. The effect on total production cost will depend on prompts, thinking level, tool use, and task length.
  • Latency targets and rate limits: The bundle provides measured output-speed results and API capabilities, but it does not establish universal latency targets or specific rate limits for every access route.
  • Open-source status: The supplied material does not identify model weights or an open-source license. Artificial Analysis labels Gemini 3.8 Flash a proprietary model.
  • Long-context behavior: The 1,048,576-token input limit is confirmed, but broader independent testing is still needed to establish how quality changes across different portions of that context window.

The model's existence, launch date, price, model ID, context limit, modalities, and primary access routes are no longer open questions. The remaining uncertainty concerns how consistently the released system performs across real workloads and how its additional reasoning affects production economics.

Frequently Asked Questions

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google's generally available Flash model, launched on September 2, 2026. Its model ID is gemini-3.8-flash, and it is designed for long-horizon software engineering, autonomous agents, and complex enterprise workflows.

Is Gemini 3.8 Flash released?

Yes. Google launched Gemini 3.8 Flash on September 2, 2026, and made it generally available through the Gemini API and Google AI Studio. It is also available in the Gemini app for Google AI Pro and Ultra users, as well as through several developer platforms.

When will Gemini 3.8 Flash be released?

Gemini 3.8 Flash launched on September 2, 2026. Google announced the model as generally available, with access through the Gemini API and Google AI Studio.

How much does Gemini 3.8 Flash cost?

Gemini 3.8 Flash costs $0.75 per 1 million input tokens and $3.75 per 1 million output tokens at its introductory price. The supplied launch reporting states that this introductory pricing runs through the end of 2026.

Does Gemini 3.8 Flash have a 1M-token context window?

Yes. Google's Gemini API documentation lists an input token limit of 1,048,576, or approximately 1 million tokens, and a maximum output token limit of 65,536.

Gemini 3.8 Flash vs Gemini 3.7 Flash: what is different?

Gemini 3.8 Flash is the newer generally available model. Google describes significant improvements over 3.7 Flash in software engineering, agentic tasks, and critical multi-step reasoning. It can use more tokens on complex tasks because it takes additional reasoning steps and verifies its work more often.

Is Gemini 3.8 Flash open source?

The supplied launch and documentation material does not identify open weights or an open-source license for Gemini 3.8 Flash. Artificial Analysis labels it a proprietary model.

What to Watch Next

The next meaningful signals are broader independent evaluations of coding, tool execution, long-context reliability, latency, and total cost per task. The released model already has official specifications and multiple access routes, so the central question has shifted from whether Gemini 3.8 Flash will launch to how well its extra reasoning and verification hold up across production workloads.

Building similar cost-focused, long-context AI workflows? On kie.ai you can try Gemini 3.8 Flash, DeepSeek-V4.1-Flash, and Claude Sonnet 5.

Maya Chen

About Maya Chen

Maya tracks AI model releases, benchmarks, and developer adoption signals across the open and closed model landscape.

View all posts by Maya Chen