What Is Qwen3.8-2.4T-A95B? Alibaba's 2.4T MoE

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: August 17, 2026
Qwen3.8-2.4T-A95B open-weight mixture-of-experts model from Alibaba

TLDRQwen3.8-2.4T-A95B is Alibaba's open-weight MoE: 2.4T total params, 95B active, 512 experts, up to 1M context, priced $2/$6 per 1M tokens.

Qwen3.8-2.4T-A95B 101: Alibaba's 2.4T Open-Weight MoE With 95B Active Params

Qwen3.8-2.4T-A95B is Alibaba's open-weight release of its Qwen3.8-Max flagship, a sparse mixture-of-experts language model with 2.4 trillion total parameters and 95 billion active per token. The weights landed on Hugging Face on August 12, 2026, routed across 512 experts under a custom qwen3.8-max license. It targets coding, research, and long-horizon agentic work, with hosted API pricing around $2.00 per million input tokens and $6.00 per million output.

Key Takeaways

  • Qwen3.8-2.4T-A95B is the open-weight version of Qwen3.8-Max, Alibaba's largest open model to date, released August 12, 2026.
  • It uses a fine-grained MoE architecture: 2.4 trillion total parameters, 95 billion active per token, 512 experts.
  • The design mixes full-attention and linear-attention layers to keep compute and KV-cache memory bounded as context scales toward 1 million tokens.
  • Hosted pricing is roughly $2.00 / $6.00 per million input/output tokens across early providers.
  • The open weights are text-only; community members flagged that vision and the full 1M context tied to hosted Qwen3.8-Max are absent from the download.
  • Serving needs data-center hardware, but quantized builds down to 397 GB (Unsloth Dynamic 1-bit) enable constrained local runs.

What Is Qwen3.8-2.4T-A95B?

Qwen3.8-2.4T-A95B is a text-generation large language model developed by Alibaba's Qwen team. It is the open-weight distribution of Qwen3.8-Max, the flagship that Alibaba previewed at the World AI Conference in Shanghai and made generally available on August 3, 2026. The open weights followed on August 12, 2026, published on Hugging Face and ModelScope.

The name encodes the architecture. "2.4T" is the total parameter count, 2.4 trillion. "A95B" is the active parameter count, 95 billion per token. That gap between total and active is the defining trait of a sparse mixture-of-experts model: capacity is enormous, but only a fraction fires on any given token.

Alibaba positions it for demanding reasoning and agentic workloads. The model is confirmed released, not leaked. On Hugging Face it carries more than 1,000 likes and is tagged as a qwen3_5_moe_text conversational model. One point of friction: the download is text-only, a fact several community members raised directly on the model card after expecting the full multimodal Qwen3.8-Max feature set.

Qwen3.8-2.4T-A95B at a Glance

AttributeDetail
DeveloperAlibaba (Qwen team)
TypeSparse mixture-of-experts (MoE) LLM
Total parameters2.4 trillion
Active parameters95 billion per token
Experts512 (fine-grained MoE)
AttentionHybrid full + linear (Gated DeltaNet / Gated Attention)
ModalityText-only (open-weight release)
Context window262K native, provider claims up to 1M (unconfirmed single spec)
Max outputUp to ~128K tokens
Hosted pricing~$2.00 in / $6.00 out per 1M tokens (provider-reported)
AvailabilityOpen weights on Hugging Face + ModelScope; hosted on multiple providers
LicenseCustom qwen3.8-max license (27B sibling is Apache 2.0)
Release dateAugust 12, 2026 (open weights)

How Qwen3.8-2.4T-A95B Works

The model is a Fine-Grained MoE. Instead of a few large experts, capacity is spread across 512 smaller experts, and a learned router activates only the experts a token needs. Serving cost tracks the 95 billion active parameters, not the full 2.4 trillion, which is what makes a model this size practical to run.

Attention is the second differentiator. Qwen3.8-2.4T-A95B alternates between full-attention and linear-attention layers in a Hybrid Attention design. In full-attention layers, every token attends to every other token. In linear-attention layers, the growing KV cache is replaced with a bounded recurrent state. NVIDIA's engineering writeup describes this as the mechanism that keeps "both compute and memory bounded as context scales to up to one million tokens," per NVIDIA's technical blog.

A third feature is configurable reasoning. Built-in Reasoning Controls (low / high / xhigh) let developers set inference depth per request, trading compute for reasoning quality. Dial up for multi-step reasoning, dial down for high-throughput document processing.

On raw serving numbers, NVIDIA reported over 4,000 tokens per second per GPU and over 350 tokens per second per user on a GB300 NVL72 rack in FP8 precision on day zero, before further tuning. That is a hardware-specific figure from NVIDIA, not an independent benchmark.

What You Can Do With Qwen3.8-2.4T-A95B

The model is built for long-context, tool-heavy work. Agentic applications accumulate system instructions, tool outputs, retrieved documents, and multi-step reasoning traces, and the hybrid attention design is meant to hold up as that context grows.

Reported use cases from the launch material include autonomous coding, large-scale document analysis, and long-running multi-step workflows. ModelScope claimed the model coded unattended for more than 10 days, reproduced a research paper, ran a 125-hour autonomous loop, and reduced a chip die area by 81% from RTL to layout. Those are platform claims without published methodology, so treat them as promotional rather than verified.

Independent throughput testing has started. Fixstars reported roughly twice the throughput of Kimi K3 under the same conditions on 8× B300 GPUs, with zero failures across 48 concurrent requests, in a day-one deployment test. If you want to compare against Moonshot's frontier model directly, see our Kimi K3 release breakdown. For teams that prefer a hosted chat endpoint over standing up 2.4T-parameter infrastructure, a comparable frontier model like Kimi K3 is available through a simple API.

How Qwen3.8-2.4T-A95B Compares

Qwen3.8-2.4T-A95B competes at the open-weight frontier alongside Moonshot's Kimi K3 and other 2T-class models. On price, a third-party model directory lists it well below Claude Opus 5.

ModelInput / Output (per 1M tokens)ContextWeights
Qwen3.8-2.4T-A95B~$2.00 / $6.00up to 1M (claims vary)Open
Kimi K3$3.00 / $15.001MOpen
Claude Opus 5$5.00 / $25.001MClosed

Pricing figures come from provider listings and a third-party directory; they are hosted rates, not official Alibaba rate cards for the open weights.

Availability: How to Access Qwen3.8-2.4T-A95B

The primary channel is the official Qwen3.8-2.4T-A95B model repository on Hugging Face, which ships weights and config for use with vLLM, SGLang, and Transformers. Alibaba also published the weights to ModelScope and its own Qwen Cloud and Model Studio channels.

For hosted inference, Alibaba promoted several day-0 launch partners on its official account, and multiple providers announced live support in the days after release. NVIDIA published deployment guidance for the model on GB300 NVL72 hardware through NIM, noting that Kubernetes is the only supported method and that a minimum of four nodes is required.

For constrained local runs, Unsloth published a quantized build. According to Unsloth's announcement, Dynamic 1-bit quantization shrank the model from 4.9 TB to 397 GB, a 91% reduction, runnable on 410 GB or more of combined RAM and VRAM. Qwen3.8 can now be run locally! 🔥 We shrank Qwen3.8-2.4T-A95B from 4.9TB to 397GB (-91% size) via D

Source: @UnslothAI

What We Don't Know Yet

Several specifics remained unsettled at the time of writing:

  • Context window. Provider claims span 256K, native 262K, and a full 1M tokens, without a single authoritative figure clarifying whether 1M is native or extended.
  • License terms. The 27B sibling is Apache 2.0, but the exact license for the 2.4T model is a custom qwen3.8-max license whose terms were not spelled out in the source material.
  • Benchmarks. Headline agentic claims (e.g. a 57% DeepSWE result at $3.73 per task) come from single commentary accounts without methodology. Independent, reproducible scores were still thin.
  • Quantized quality. Reports range from a 397 GB Dynamic 1-bit build to sub-128 GB Q1_0 operation, with some testers noting visible block artifacts at 2-bit.

Frequently Asked Questions

What is Qwen3.8-2.4T-A95B?

Qwen3.8-2.4T-A95B is Alibaba's open-weight release of its Qwen3.8-Max flagship, a sparse mixture-of-experts language model with 2.4 trillion total parameters and 95 billion active per token routed across 512 experts. It was released on Hugging Face on August 12, 2026 under a custom qwen3.8-max license.

How many parameters does Qwen3.8-2.4T-A95B have?

Qwen3.8-2.4T-A95B has 2.4 trillion total parameters with 95 billion active per token, which is what the "A95B" in the name denotes. The sparse MoE design routes each token to a subset of 512 experts, so serving cost tracks the 95B active count rather than the full 2.4T.

Is Qwen3.8-2.4T-A95B open source?

Qwen3.8-2.4T-A95B is open-weight, with model weights and config files published on Hugging Face on August 12, 2026. The license shown on the model card is a custom "qwen3.8-max" license; the sibling Qwen3.8-27B model ships under Apache 2.0, but the 2.4T model's exact license terms were not confirmed as Apache 2.0 in the source material.

How much does Qwen3.8-2.4T-A95B cost?

Hosted API pricing reported by inference providers is around $2.00 per million input tokens and $6.00 per million output tokens, with cached input near $0.20 per million. Because the weights are open, self-hosting is also possible, though it requires data-center-scale hardware.

What is the context window of Qwen3.8-2.4T-A95B?

Context window figures vary by provider, ranging from a native 262K tokens up to 1 million tokens. Provider claims in the source material span 256K, native 262K, and a full 1M-token context; the authoritative single specification was not settled at the time of writing.

Can I run Qwen3.8-2.4T-A95B locally?

Running Qwen3.8-2.4T-A95B locally is possible but hardware-intensive. According to Unsloth, a Dynamic 1-bit quantization shrank the model from 4.9 TB to 397 GB and requires 410 GB or more of combined RAM and VRAM; other testers report sub-128 GB Q1_0 builds with visible quality degradation.

Is Qwen3.8-2.4T-A95B multimodal?

The open-weight Qwen3.8-2.4T-A95B release is text-only. Community members noted the released weights are stripped of the vision and 1M-context features associated with the hosted Qwen3.8-Max, and NVIDIA's deployment notes list image and video input as unsupported.

What to Watch Next

Three signals will shape this page's next update. First, an authoritative context-window number: whether Alibaba confirms 262K native or a true 1M window resolves the biggest open spec question. Second, independent benchmarks with published methodology, which will test the "Max-level" and "rivals GPT-5.6 Sol" claims that currently rest on provider marketing. Third, license clarity on the 2.4T weights, since the gap between the 27B model's Apache 2.0 and the custom qwen3.8-max license matters for commercial adopters.

Building similar large-scale, long-context chat and agentic workflows? On kie.ai you can try Gemini 3.8 Flash, Claude Opus 5, and Grok 4.6.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair