What Is Qwen3.8-2.4T-A95B? Alibaba's 2.4T MoE
Priya Nair
AI Infrastructure Analyst

TLDRQwen3.8-2.4T-A95B is Alibaba's open-weight MoE: 2.4T total params, 95B active, 512 experts, up to 1M context, priced $2/$6 per 1M tokens.
Qwen3.8-2.4T-A95B 101: Alibaba's 2.4T Open-Weight MoE With 95B Active Params
Qwen3.8-2.4T-A95B is Alibaba's open-weight release of its Qwen3.8-Max flagship, a sparse mixture-of-experts language model with 2.4 trillion total parameters and 95 billion active per token. The weights landed on Hugging Face on August 12, 2026, routed across 512 experts under a custom qwen3.8-max license. It targets coding, research, and long-horizon agentic work, with hosted API pricing around $2.00 per million input tokens and $6.00 per million output.
Key Takeaways
- Qwen3.8-2.4T-A95B is the open-weight version of Qwen3.8-Max, Alibaba's largest open model to date, released August 12, 2026.
- It uses a fine-grained MoE architecture: 2.4 trillion total parameters, 95 billion active per token, 512 experts.
- The design mixes full-attention and linear-attention layers to keep compute and KV-cache memory bounded as context scales toward 1 million tokens.
- Hosted pricing is roughly $2.00 / $6.00 per million input/output tokens across early providers.
- The open weights are text-only; community members flagged that vision and the full 1M context tied to hosted Qwen3.8-Max are absent from the download.
- Serving needs data-center hardware, but quantized builds down to 397 GB (Unsloth Dynamic 1-bit) enable constrained local runs.
What Is Qwen3.8-2.4T-A95B?
Qwen3.8-2.4T-A95B is a text-generation large language model developed by Alibaba's Qwen team. It is the open-weight distribution of Qwen3.8-Max, the flagship that Alibaba previewed at the World AI Conference in Shanghai and made generally available on August 3, 2026. The open weights followed on August 12, 2026, published on Hugging Face and ModelScope.
The name encodes the architecture. "2.4T" is the total parameter count, 2.4 trillion. "A95B" is the active parameter count, 95 billion per token. That gap between total and active is the defining trait of a sparse mixture-of-experts model: capacity is enormous, but only a fraction fires on any given token.
Alibaba positions it for demanding reasoning and agentic workloads. The model is confirmed released, not leaked. On Hugging Face it carries more than 1,000 likes and is tagged as a qwen3_5_moe_text conversational model. One point of friction: the download is text-only, a fact several community members raised directly on the model card after expecting the full multimodal Qwen3.8-Max feature set.
Qwen3.8-2.4T-A95B at a Glance
| Attribute | Detail |
|---|---|
| Developer | Alibaba (Qwen team) |
| Type | Sparse mixture-of-experts (MoE) LLM |
| Total parameters | 2.4 trillion |
| Active parameters | 95 billion per token |
| Experts | 512 (fine-grained MoE) |
| Attention | Hybrid full + linear (Gated DeltaNet / Gated Attention) |
| Modality | Text-only (open-weight release) |
| Context window | 262K native, provider claims up to 1M (unconfirmed single spec) |
| Max output | Up to ~128K tokens |
| Hosted pricing | ~$2.00 in / $6.00 out per 1M tokens (provider-reported) |
| Availability | Open weights on Hugging Face + ModelScope; hosted on multiple providers |
| License | Custom qwen3.8-max license (27B sibling is Apache 2.0) |
| Release date | August 12, 2026 (open weights) |
How Qwen3.8-2.4T-A95B Works
The model is a Fine-Grained MoE. Instead of a few large experts, capacity is spread across 512 smaller experts, and a learned router activates only the experts a token needs. Serving cost tracks the 95 billion active parameters, not the full 2.4 trillion, which is what makes a model this size practical to run.
Attention is the second differentiator. Qwen3.8-2.4T-A95B alternates between full-attention and linear-attention layers in a Hybrid Attention design. In full-attention layers, every token attends to every other token. In linear-attention layers, the growing KV cache is replaced with a bounded recurrent state. NVIDIA's engineering writeup describes this as the mechanism that keeps "both compute and memory bounded as context scales to up to one million tokens," per NVIDIA's technical blog.
A third feature is configurable reasoning. Built-in Reasoning Controls (low / high / xhigh) let developers set inference depth per request, trading compute for reasoning quality. Dial up for multi-step reasoning, dial down for high-throughput document processing.
On raw serving numbers, NVIDIA reported over 4,000 tokens per second per GPU and over 350 tokens per second per user on a GB300 NVL72 rack in FP8 precision on day zero, before further tuning. That is a hardware-specific figure from NVIDIA, not an independent benchmark.
What You Can Do With Qwen3.8-2.4T-A95B
The model is built for long-context, tool-heavy work. Agentic applications accumulate system instructions, tool outputs, retrieved documents, and multi-step reasoning traces, and the hybrid attention design is meant to hold up as that context grows.
Reported use cases from the launch material include autonomous coding, large-scale document analysis, and long-running multi-step workflows. ModelScope claimed the model coded unattended for more than 10 days, reproduced a research paper, ran a 125-hour autonomous loop, and reduced a chip die area by 81% from RTL to layout. Those are platform claims without published methodology, so treat them as promotional rather than verified.
Independent throughput testing has started. Fixstars reported roughly twice the throughput of Kimi K3 under the same conditions on 8× B300 GPUs, with zero failures across 48 concurrent requests, in a day-one deployment test. If you want to compare against Moonshot's frontier model directly, see our Kimi K3 release breakdown. For teams that prefer a hosted chat endpoint over standing up 2.4T-parameter infrastructure, a comparable frontier model like Kimi K3 is available through a simple API.
How Qwen3.8-2.4T-A95B Compares
Qwen3.8-2.4T-A95B competes at the open-weight frontier alongside Moonshot's Kimi K3 and other 2T-class models. On price, a third-party model directory lists it well below Claude Opus 5.
| Model | Input / Output (per 1M tokens) | Context | Weights |
|---|---|---|---|
| Qwen3.8-2.4T-A95B | ~$2.00 / $6.00 | up to 1M (claims vary) | Open |
| Kimi K3 | $3.00 / $15.00 | 1M | Open |
| Claude Opus 5 | $5.00 / $25.00 | 1M | Closed |
Pricing figures come from provider listings and a third-party directory; they are hosted rates, not official Alibaba rate cards for the open weights.
Availability: How to Access Qwen3.8-2.4T-A95B
The primary channel is the official Qwen3.8-2.4T-A95B model repository on Hugging Face, which ships weights and config for use with vLLM, SGLang, and Transformers. Alibaba also published the weights to ModelScope and its own Qwen Cloud and Model Studio channels.
For hosted inference, Alibaba promoted several day-0 launch partners on its official account, and multiple providers announced live support in the days after release. NVIDIA published deployment guidance for the model on GB300 NVL72 hardware through NIM, noting that Kubernetes is the only supported method and that a minimum of four nodes is required.
For constrained local runs, Unsloth published a quantized build. According to Unsloth's announcement, Dynamic 1-bit quantization shrank the model from 4.9 TB to 397 GB, a 91% reduction, runnable on 410 GB or more of combined RAM and VRAM.

Source: @UnslothAI
What We Don't Know Yet
Several specifics remained unsettled at the time of writing:
- Context window. Provider claims span 256K, native 262K, and a full 1M tokens, without a single authoritative figure clarifying whether 1M is native or extended.
- License terms. The 27B sibling is Apache 2.0, but the exact license for the 2.4T model is a custom
qwen3.8-maxlicense whose terms were not spelled out in the source material. - Benchmarks. Headline agentic claims (e.g. a 57% DeepSWE result at $3.73 per task) come from single commentary accounts without methodology. Independent, reproducible scores were still thin.
- Quantized quality. Reports range from a 397 GB Dynamic 1-bit build to sub-128 GB Q1_0 operation, with some testers noting visible block artifacts at 2-bit.
Frequently Asked Questions
What is Qwen3.8-2.4T-A95B?
Qwen3.8-2.4T-A95B is Alibaba's open-weight release of its Qwen3.8-Max flagship, a sparse mixture-of-experts language model with 2.4 trillion total parameters and 95 billion active per token routed across 512 experts. It was released on Hugging Face on August 12, 2026 under a custom qwen3.8-max license.
How many parameters does Qwen3.8-2.4T-A95B have?
Qwen3.8-2.4T-A95B has 2.4 trillion total parameters with 95 billion active per token, which is what the "A95B" in the name denotes. The sparse MoE design routes each token to a subset of 512 experts, so serving cost tracks the 95B active count rather than the full 2.4T.
Is Qwen3.8-2.4T-A95B open source?
Qwen3.8-2.4T-A95B is open-weight, with model weights and config files published on Hugging Face on August 12, 2026. The license shown on the model card is a custom "qwen3.8-max" license; the sibling Qwen3.8-27B model ships under Apache 2.0, but the 2.4T model's exact license terms were not confirmed as Apache 2.0 in the source material.
How much does Qwen3.8-2.4T-A95B cost?
Hosted API pricing reported by inference providers is around $2.00 per million input tokens and $6.00 per million output tokens, with cached input near $0.20 per million. Because the weights are open, self-hosting is also possible, though it requires data-center-scale hardware.
What is the context window of Qwen3.8-2.4T-A95B?
Context window figures vary by provider, ranging from a native 262K tokens up to 1 million tokens. Provider claims in the source material span 256K, native 262K, and a full 1M-token context; the authoritative single specification was not settled at the time of writing.
Can I run Qwen3.8-2.4T-A95B locally?
Running Qwen3.8-2.4T-A95B locally is possible but hardware-intensive. According to Unsloth, a Dynamic 1-bit quantization shrank the model from 4.9 TB to 397 GB and requires 410 GB or more of combined RAM and VRAM; other testers report sub-128 GB Q1_0 builds with visible quality degradation.
Is Qwen3.8-2.4T-A95B multimodal?
The open-weight Qwen3.8-2.4T-A95B release is text-only. Community members noted the released weights are stripped of the vision and 1M-context features associated with the hosted Qwen3.8-Max, and NVIDIA's deployment notes list image and video input as unsupported.
What to Watch Next
Three signals will shape this page's next update. First, an authoritative context-window number: whether Alibaba confirms 262K native or a true 1M window resolves the biggest open spec question. Second, independent benchmarks with published methodology, which will test the "Max-level" and "rivals GPT-5.6 Sol" claims that currently rest on provider marketing. Third, license clarity on the 2.4T weights, since the gap between the 27B model's Apache 2.0 and the custom qwen3.8-max license matters for commercial adopters.
Building similar large-scale, long-context chat and agentic workflows? On kie.ai you can try Gemini 3.8 Flash, Claude Opus 5, and Grok 4.6.
About Priya Nair
Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.
View all posts by Priya Nair