What Is Qwen3.8-35B? 3B-Active Model

Daniel Okonkwo

Daniel Okonkwo

Senior ML Engineer

Published: September 1, 2026
Qwen3.8-35B model overview

TLDRQwen3.8-35B is a proposed 35B-A3B MoE model with no confirmed checkpoint, price, context limit, or official access path.

Inside Qwen3.8-35B: The 3B-Active Model Local Users Keep Requesting

Qwen3.8-35B is an unreleased, unconfirmed Qwen model name for a proposed 35B-parameter mixture-of-experts model with roughly 3B active parameters per token. It is associated with Alibaba’s Qwen family, but no official Qwen3.8-35B checkpoint, model card, benchmark, price, or release announcement has been confirmed. The name first surfaced through early repository and tooling reports, while community demand has focused on a faster local model between Qwen3.8-27B and larger Qwen variants.

Key Takeaways

  • Qwen3.8-35B is currently a proposed model identity, not a confirmed downloadable release.
  • The commonly requested configuration is 35B-A3B, meaning 35B total parameters and approximately 3B active parameters per token; that specification remains unconfirmed.
  • A reported ModelScope repository appearance was followed by a deletion report, creating the Repository Signal around the model name.
  • A later ms-swift correction reportedly replaced Qwen3.8-35B identifiers with Qwen3.8-27B identifiers. This is the Model-ID Correction.
  • Local users want the model for higher interactive throughput than Qwen3.8-27B, especially on systems with 8GB to 32GB of VRAM.
  • Qwen3.8-35B has no confirmed context window, license, API, download location, or official pricing.

What Is Qwen3.8-35B?

Qwen3.8-35B is the name used by users discussing a possible Qwen3.8-family 35B-A3B model. In the requested configuration, “35B” refers to the model’s total parameter class, while “A3B” refers to the approximate number of parameters activated for each token. This is a community interpretation of the name, not a published Qwen specification.

The model’s status is the central fact. On August 18, 2026, a third-party account reported that Qwen3.8-35B-A3B appeared in a ModelScope-related repository commit, then reported that the entry had been deleted. The report described the appearance as a possible accidental leak. The repository signal can be reviewed through the account’s original post. The initial ModelScope report

On August 19, 2026, another report said an ms-swift support change had initially included Qwen3.8-35B-A3B and Qwen3.8-35B-A3B-FP8, before a later commit titled “fix wrong model-ids” replaced them with Qwen3.8-27B identifiers. The reported model-ID correction The linked GitHub commit is a more concrete artifact than ordinary speculation, but it still does not prove that a public 35B checkpoint exists. The referenced ms-swift commit

Qwen3.8-35B is not a released model; it is an unconfirmed name for a requested Qwen mixture-of-experts variant.

The confirmed Qwen3.8 material in the supplied evidence points to Qwen3.8-27B and a 2.4T-A95B model, not to a confirmed 35B-A3B release. A report summarizing the lineup said the official list contained those two models and not 35B-A3B. The lineup and correction report

Qwen3.8-35B at a Glance

SpecificationCurrent status
DeveloperQwen team associated with Alibaba; 35B attribution not formally confirmed
TypeProposed mixture-of-experts language model
Parameter profile35B total / approximately 3B active, based on community shorthand; unconfirmed
ModalityNot yet confirmed
Context windowNot yet confirmed
PricingNot yet confirmed
AvailabilityNo confirmed official checkpoint or download
LicenseNot yet confirmed

For comparison, the confirmed Qwen3.8-27B model page identifies an image-text-to-text model and lists the Apache-2.0 license. Those facts belong to Qwen3.8-27B and should not be transferred to Qwen3.8-35B without a separate model card. Qwen3.8-27B on Hugging Face

How Qwen3.8-35B Works / What Makes It Different

The expected design is a 35B-A3B Profile: a mixture-of-experts model with a large total parameter pool but a smaller active path for each token. In an MoE system, routing selects a subset of experts rather than evaluating every parameter for every token. If the community’s A3B description is accurate, only about 3B parameters would be active during each token calculation, although the full model would still need to be stored or streamed.

A3B describes the community’s expected active-parameter target, not a confirmed Qwen3.8 specification.

That distinction explains the model’s appeal. A dense 27B model can deliver strong quality but may process tokens more slowly on unified-memory machines. A 35B-A3B model could, in principle, combine a larger capacity pool with a lower per-token compute path. The actual speed would depend on routing, memory bandwidth, quantization, context length, and inference software. No Qwen3.8-35B test verifies the expected trade-off.

The Local Hardware Gap is the practical problem behind most requests. Community posts describe Qwen3.8-27B as capable but slow for long reasoning tasks, while users report much higher throughput from older 35B-A3B models. One Reddit user reported approximately 120 tokens per second with an older 35B-A3B model versus approximately 20 tokens per second with Qwen3.8-27B. That comparison is a personal test, not a Qwen3.8-35B result. The local inference discussion

The proposed model would therefore target interactive local inference. Users are asking for a model that can preserve more capability than smaller local models while avoiding the waiting time associated with extended reasoning on Qwen3.8-27B. Whether Qwen3.8-35B would achieve that goal remains unknown.

What You Can Do With Qwen3.8-35B

There is currently nothing reliable to run. No confirmed checkpoint or official endpoint is available in the supplied evidence.

The intended use cases are nevertheless clear from the community discussion:

  • Interactive local coding: Users want shorter waits during software development and agentic coding sessions.
  • Tool-using agents: Posts discuss Hermes Agents, tool-call reliability, and the desire for a faster model that can operate interactively.
  • Consumer-GPU inference: The requested A3B profile is aimed at systems with limited VRAM, including 8GB GPUs and machines with 16GB or 32GB of VRAM.
  • Long-context local work: Community comparisons discuss context windows of 64K, 128K, 200K, and 262K tokens on related models. None of those figures is confirmed for Qwen3.8-35B.
  • Overnight or batch tasks: Users specifically contrast the time required by Qwen3.8-27B with the hoped-for throughput of a 35B-A3B model.

The Proxy Fit Test should be kept separate from direct evidence. A community test ran Ornith-1.5-35B-A3B-NVFP4 at 22 tokens per second on an RTX 4060 with 8GB of VRAM and 64GB of system RAM, using a 64K-token context and 7.5GB of actual VRAM. The author inferred that a similarly quantized Qwen3.8-35B might fit, but did not test Qwen3.8-35B. The proxy hardware report

How Qwen3.8-35B Compares

The closest confirmed reference point is Qwen3.8-27B, while Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B appear in community discussions as practical alternatives.

ModelStatusReported shape or result
Qwen3.8-35BUnconfirmedProposed 35B-A3B MoE profile; no direct test
Qwen3.8-27BConfirmed checkpoint27B model page available; Apache-2.0 listed
Qwen3.6-35B-A3BEarlier related modelReported at 50 tokens per second with 262K-token context on an RTX 5080 with 16GB VRAM and system-memory offload
Ornith-1.5-35B-A3BCommunity stopgapReported at 22 tokens per second with 64K-token context on an RTX 4060 with 8GB VRAM

A reported 50-token-per-second result for Qwen3.6-35B-A3B does not establish Qwen3.8-35B performance. The Qwen3.6 community benchmark

For background on the confirmed 27B model, see the earlier analysis of Qwen 3.8 27B’s release and dense-model positioning and the first look at Qwen 3.8 27B for coding workloads.

Availability: How to Access Qwen3.8-35B

Qwen3.8-35B has no confirmed official access path. There is no verified Qwen model page, API endpoint, application listing, download archive, or published inference guide for this specific name.

The safest verification path is to check Qwen’s official release announcements, model repositories, and documentation for a new checkpoint. A repository remnant, an inference-library identifier, or a community distillation should not be treated as the official model. In particular, reports of hobbyist “Qwen3.8” distillations describe them as derived from other bases, so they cannot establish the existence of an official Qwen3.8-35B model.

Kie.ai does not currently host Qwen3.8-35B. For a hosted chat model while the status remains unresolved, Gemini 3.7 Flash is a separate catalog model and is not a Qwen3.8-35B substitute or release channel.

What We Don't Know Yet

The unresolved questions are more important than the leaked identifier:

  • Whether Alibaba’s Qwen team ever intended to release a 35B-A3B model.
  • Whether the reported repository entry was a real checkpoint reference or an accidental configuration error.
  • Whether the A3B suffix describes approximately 3B active parameters in an official architecture.
  • Whether Qwen3.8-35B would be text-only, multimodal, or support image input like the confirmed Qwen3.8-27B page.
  • What context window, quantization formats, VRAM requirement, and inference frameworks it would support.
  • Whether Qwen3.8-Flash-Next is related to, derived from, or simply a replacement for the requested 35B-A3B model.

The Flash-Next Question remains open. A report said Qwen3.8-Flash-Next would become downloadable on August 27, 2026, but did not establish equivalence with Qwen3.8-35B-A3B. The Flash-Next comparison question

No Qwen3.8-35B checkpoint, official context limit, price, or benchmark is confirmed as of September 1, 2026.

Frequently Asked Questions

Is Qwen3.8-35B released?

Qwen3.8-35B is not confirmed as released. Reports of a repository entry and tooling identifiers were followed by deletion and a correction that replaced the identifiers with Qwen3.8-27B identifiers.

What is Qwen3.8-35B A3B?

Qwen3.8-35B A3B is the community name for a proposed mixture-of-experts model with 35B total parameters and approximately 3B active parameters per token. Qwen has not confirmed those specifications for an official Qwen3.8 checkpoint.

Is Qwen3.8-35B open source?

Qwen3.8-35B is not currently confirmed as open source because no official checkpoint or license has been published for it. The confirmed Qwen3.8-27B model page lists Apache-2.0, but that license does not establish the license for a 35B variant.

How much does Qwen3.8-35B cost?

The cost of Qwen3.8-35B is not yet confirmed. No official API price, hosted-product price, or download fee is available in the supplied release evidence.

How can I download Qwen3.8-35B?

Qwen3.8-35B cannot currently be downloaded from a confirmed official model page. Builders should verify the official Qwen channels for a checkpoint rather than treating community distillations, repository remnants, or similarly named models as an official release.

Qwen3.8-35B vs Qwen3.8-27B?

Qwen3.8-35B is an unconfirmed proposed 35B-A3B MoE variant, while Qwen3.8-27B is a confirmed 27B checkpoint with an available model page. The 35B model has no verified benchmark or quality comparison against the 27B model.

What to Watch Next

The most useful update signals are an official Qwen model card, a downloadable checkpoint with a license, and a verified benchmark using the 35B-A3B identifier. A clear statement about Qwen3.8-Flash-Next would also resolve whether it is related to the requested model or merely another branch of the lineup.

Building similar efficient local-first chat and coding workflows? On kie.ai you can try DeepSeek-V4.1-Flash, Gemini 3.8 Flash, and Claude Sonnet 5.

Daniel Okonkwo

About Daniel Okonkwo

Daniel writes about inference systems, model architecture, and what new releases actually change for builders.

View all posts by Daniel Okonkwo