What Is Qwen3.8 Flash: 1M Context AI Model
Daniel Okonkwo
Senior ML Engineer

TLDRA 125B/6B-active multimodal MoE with 1M-token API context and $0.15/$0.47 per-million-token pricing.
1M-Token Context: A First Look at Qwen3.8 Flash
Qwen3.8 Flash is Alibaba Qwen’s production multimodal mixture-of-experts model for reasoning, coding, tool use, and visual understanding. Its hosted API supports a 1M-token context window and lists pricing of $0.15 per 1M input tokens and $0.47 per 1M output tokens. The related Qwen3.8-Flash-Next checkpoint is an open-weight preview of the architecture intended to underpin Qwen4. The production model is available through Qwen’s official cloud channels, while the preview weights are published for local experimentation.
Key Takeaways
- Qwen3.8 Flash is a multimodal MoE model with text, image, and video input and text output.
- Its production API provides a 1M-token context window, with a maximum output of 131K tokens listed by QwenCloud.
- The related open-weight Flash-Next model has 125B main parameters, 6B active parameters per token, and an additional 51B N-gram Embedding parameter set.
- Qwen3.8-Flash-Next uses Gated DeltaNet, Qwen Sparse Attention, and Gated Residual to reduce long-context compute and improve information flow.
- QwenCloud lists input pricing at $0.15 per 1M tokens and output pricing at $0.47 per 1M tokens.
- Published use cases include agentic coding, document processing, tool-driven workflows, chart analysis, video understanding, and computer-use tasks.
What Is Qwen3.8 Flash?
Qwen3.8 Flash is the fast, cost-oriented production member of Alibaba’s Qwen model family. It combines language reasoning with native image and video understanding, so one API can handle software tasks, documents, charts, screenshots, and long videos alongside ordinary text.
The model is closely related to Qwen3.8-Flash-Next. Qwen describes Flash-Next as an early architectural release that lets developers examine the design before the broader Qwen4 family. The official Qwen architecture announcement says the preview opens the weights and serves as an early view of Qwen4.
This distinction matters. “Qwen3.8 Flash” generally refers to the production API model, while “Qwen3.8-Flash-Next” refers to the open-weight preview checkpoint. The two names describe connected but different distribution forms.
Qwen3.8 Flash is best understood as a production API built around the Flash-Next architectural preview.
The Flash positioning is about efficient inference rather than maximum total parameter count. The open-weight preview contains 125B main-model parameters, but only 6B are activated for each token. It also adds 51B parameters in an N-gram Embedding table. This sparse design aims to deliver stronger capability than a small dense model without computing every parameter on every token.
Qwen’s announcement also states that training the preview cost about one-ninth as much as Qwen3.7-Plus. That figure is a vendor-published comparison, not an independently audited training-cost estimate.
Qwen3.8 Flash at a Glance
| Specification | Qwen3.8 Flash |
|---|---|
| Developer | Alibaba Qwen |
| Type | Multimodal mixture-of-experts model |
| Modality | Text, image, and video input; text output |
| Main model parameters | 125B in the related Flash-Next preview |
| Active parameters | 6B per token in the related Flash-Next preview |
| Additional embedding parameters | 51B N-gram Embedding parameters |
| Context window | 1M tokens in the production API |
| Native preview context | 262,144 tokens, extensible to 1,000,000 tokens with YaRN |
| Maximum input | 991K tokens listed by QwenCloud |
| Maximum output | 131K tokens listed by QwenCloud |
| Pricing | $0.15 per 1M input tokens; $0.47 per 1M output tokens |
| Cached input pricing | $0.016 per 1M tokens listed for implicit cache reads |
| Availability | QwenCloud, Model Studio, gmi_cloud, and OpenCode Go |
| Open-weight checkpoint | Qwen3.8-Flash-Next |
| License | Qwen Community 1.0 for Flash-Next weights; production API terms not yet confirmed |
The production API listing also reports a 2M-token-per-minute limit and a 15K-requests-per-minute limit. Those limits can vary by account, region, or service configuration, so they should be checked before deployment.
How Qwen3.8 Flash Works and What Makes It Different
The defining design choices are concentrated in four architectural areas: attention, residual connections, embeddings, and optimization.
Gated DeltaNet and Qwen Sparse Attention
Gated DeltaNet compresses historical context into a recurrent state. In the published hybrid layout, three out of every four layers use this linear-attention approach. That limits the growth of key-value cache pressure as a sequence becomes longer.
The remaining attention layer uses Qwen Sparse Attention. QSA works at micro-block granularity rather than selecting individual tokens. It estimates which blocks matter, then focuses precise attention on those regions. This design targets the cost bottlenecks that appear in 262K-token and 1M-token workloads.
NVIDIA’s technical discussion reports attention-kernel speedups of up to 7.6 times during prefill and 4.9 times during decoding against full attention in its described tests. A separate 1M-token serving test with a 90% prefix-cache hit rate reported 8.6 times the prefill throughput of Qwen3.7-Plus. These are published implementation results, not a guarantee for every hardware stack.
Gated Residual
Gated Residual widens the residual stream into four branches. A data-dependent read gate controls information entering the branches, while per-branch write gates control what returns to the residual stream.
The goal is greater expressiveness without adding a large inference burden. It also addresses training stability, which becomes increasingly important as parameter capacity grows.
N-gram Embedding
N-gram Embedding adds a large lookup table indexed by local two-token and three-token patterns. The table contains 20M entries and contributes 51B additional parameters in the preview model.
Unlike ordinary dense weights, much of this table can be placed in host memory or offloaded. That creates a tradeoff: the model can scale capacity with less per-token compute, but local deployment becomes sensitive to memory bandwidth, storage access, and runtime support.
Thinking Mode
Qwen3.8-Flash-Next is a hybrid thinking model. Thinking Mode is enabled by default in the documented preview configuration, and developers can adjust reasoning effort through low, medium, or xhigh settings. Thinking can also be disabled per request.
This makes the model suitable for two different workloads. Interactive coding can use lower reasoning effort for responsiveness, while complex repository changes or multi-step agent tasks can allocate more internal reasoning.
What You Can Do With Qwen3.8 Flash
Qwen3.8 Flash targets workloads that combine long context, visual input, and external actions.
- Agentic coding: The model can inspect repositories, plan changes, call tools, and revise code across multiple steps.
- Long-document processing: A 1M-token API context can cover large document collections, long transcripts, or an entire codebase in one request.
- Tool-driven workflows: Function calling, structured outputs, web search, code execution, and context caching are listed as supported API features.
- Visual analysis: Image input supports charts, documents, diagrams, screenshots, and other visual material.
- Video understanding: The production listing describes non-real-time video understanding, including long videos.
- Computer and mobile use: The model’s published capability description includes desktop-operation and mobile-use tasks.
- Coding assistants: Qwen officially announced availability in OpenCode Go, making the model relevant to terminal-based and IDE-adjacent workflows.
The practical target is a high-volume agent that must keep substantial context while controlling inference cost.
The model’s local ecosystem also shows interest in agent applications. Ollama’s model page lists integrations with Claude Code, OpenCode, Hermes Agent, and OpenClaw for the Flash-Next checkpoint. Those are integrations around the open-weight preview, not proof that every production API feature works locally.
How Qwen3.8 Flash Compares
Qwen’s published benchmark table compares Flash-Next with Qwen3.8-27B, Qwen3.7-Plus, DeepSeek-V4-Flash-0731, and Claude Opus 4.6 Max. Selected results are below.
| Evaluation | Qwen3.8-Flash-Next | Comparison signal |
|---|---|---|
| SWE-bench Pro | 62.5 score | Claude Opus 4.6 Max: 53.4 |
| SWE-bench Multilingual | 81.0 score | Claude Opus 4.6 Max: 77.5 |
| CoWorkBench | 73.9 score | Claude Opus 4.6 Max: 68.2 |
| Toolathlon Verified | 73.5 score | DeepSeek-V4-Flash-0731: 70.3 |
| GPQA Diamond | 91.7 score | Claude Opus 4.6 Max: 91.3 |
| Humanity’s Last Exam | 35.9 score | Claude Opus 4.6 Max: 40.0 |
These numbers come from Qwen’s published evaluation results. They should not be treated as a universal capability ordering because benchmark prompts, tools, inference settings, and grading procedures affect results.
An early community comparison measured Qwen3.8 Flash at approximately 50 tokens per second on prose, compared with approximately 25 tokens per second for GLM-5.3 Flash. The same tester said pure coding speeds were much closer. The earlier analysis of GLM-5.3 community signals provides additional context on that competitor, while a separate Qwen 3.8 27B release analysis covers the smaller dense model.
Availability: How to Access Qwen3.8 Flash
The production model is available through QwenCloud and Alibaba Cloud’s Model Studio. Alibaba Cloud also announced Qwen3.8 Flash on gmi_cloud, while Qwen announced its inclusion in OpenCode Go.
QwenCloud’s current model listing identifies the production model as qwen3.8-flash, but the exact API configuration, regional endpoint, and account limits should be checked in the provider documentation. The service supports OpenAI-compatible and Anthropic-compatible protocols, according to the official listing.
For a comparable hosted chat workflow while evaluating integration patterns, Gemini 3.7 Flash is a separate chat model available in the Kie.ai catalog.
The open-weight counterpart is Qwen3.8-Flash-Next on Hugging Face. It can be served with runtimes such as vLLM, SGLang, Transformers, Ollama, llama.cpp, and MLX-based tools, although support for QSA, N-gram offloading, KV caching, and quantization differs by runtime.
The local hardware requirement is substantial. Unsloth documentation places the smallest 1-bit quantization at about 75 GB of RAM or unified memory, with higher quantizations listed at approximately 79 GB for 2-bit, 90 GB for 3-bit, 112 GB for 4-bit, and 355 GB for BF16. These figures apply to Flash-Next quantizations, not to the hosted API.
A community implementation also documents a two-DGX-Spark setup for NVFP4 inference. That dual-DGX-Spark recipe reports approximately 64 tokens per second for one stream and 117 tokens per second at two concurrent streams in its own environment. It also records runtime-specific fixes for sparse-attention kernels and KV-cache behavior.

Source: @ivanfioravanti
An early community test also demonstrated Qwen3.8-Flash-Next-MLX-Serve-4bit running on an M5 Max MacBook Pro. That result shows feasibility on high-memory consumer hardware, not a minimum official requirement.
What We Don’t Know Yet
Several details remain separate from the confirmed production specification.
- The exact relationship between the production API implementation and every component of the Flash-Next checkpoint is not fully documented.
- QwenCloud lists 1M tokens for production, while the preview weights specify 262,144 native tokens and 1,000,000 tokens with YaRN. Runtime-specific effective limits may differ.
- Independent, reproducible evaluations of the production API are still limited.
- The published benchmark scores measure Flash-Next, and they do not automatically establish identical scores for every hosted Flash version.
- Production API terms, data-retention policies, and enterprise deployment conditions are not confirmed in the supplied material.
- Local performance depends heavily on quantization, memory bandwidth, context length, speculative decoding, and N-gram offloading.
- Early DGX-Spark users reported a rare exclamation-mark output loop when using particular KV-cache paths. A later community update reported that a fix resolved the issue for multiple users, while recommending FP8 KV cache over NVFP4 in some configurations.
Frequently Asked Questions
What is Qwen3.8 Flash?
Qwen3.8 Flash is Alibaba Qwen’s production multimodal mixture-of-experts model for reasoning, coding, tool use, and vision tasks. Its API supports a 1M-token context window, while its open-weight architectural preview is named Qwen3.8-Flash-Next.
Is Qwen3.8 Flash open source?
Qwen3.8 Flash itself is documented primarily as a production API model, not as the open-weight checkpoint. Qwen3.8-Flash-Next, the related Qwen4 architecture preview, has published weights under the Qwen Community 1.0 license.
How much does Qwen3.8 Flash cost?
Qwen3.8 Flash is listed at $0.15 per 1M input tokens and $0.47 per 1M output tokens on QwenCloud. The same listing shows lower rates for cached input, subject to the provider’s current terms.
What is Qwen3.8 Flash’s context window?
Qwen3.8 Flash has a 1M-token context window in the production API. The open-weight Qwen3.8-Flash-Next preview has a native 262,144-token window and can extend to 1,000,000 tokens with YaRN.
Qwen3.8 Flash vs GLM-5.3 Flash: which is faster?
An early community comparison measured Qwen3.8 Flash at approximately 50 tokens per second versus approximately 25 tokens per second for GLM-5.3 Flash on prose generation. The same comparison said the models were much closer on pure coding, so the result is workload-specific rather than a universal speed ranking.
Can I run Qwen3.8 Flash locally?
The production Qwen3.8 Flash API is accessed through Qwen’s hosted channels, while Qwen3.8-Flash-Next can be run locally from published weights and community quantizations. Third-party documentation puts the smallest local quantization at about 75 GB of RAM or unified memory, but actual requirements depend on quantization, context length, and runtime.
What to watch next is whether Qwen publishes fuller production model documentation, whether independent tests reproduce the 1M-token efficiency claims, and how quickly runtimes stabilize N-gram offloading and sparse-attention support.
Building similar long-context coding and agent workflows? On kie.ai you can try Gemini 3.8 Flash, Claude Sonnet 4.6, and DeepSeek-V4.1-Flash.
About Daniel Okonkwo
Daniel writes about inference systems, model architecture, and what new releases actually change for builders.
View all posts by Daniel Okonkwo