Meet Xiaomi MiMo V2.6, the 1M-Token Omnimodal Model
Daniel Okonkwo
Senior ML Engineer

TLDRA 1M-token, text-image-video-audio model family from Xiaomi with open weights, RL-trained agent capabilities, and API access.
Xiaomi MiMo V2.6 is Xiaomi’s open-weight family of native omnimodal AI models for text, image, video, audio, reasoning, coding, and agent workflows. The released Pro and Flash checkpoints support a 1-million-token context window, while Xiaomi’s API lists a third Pro Ultraspeed variant for latency-sensitive workloads. MiMo V2.6 became officially available on September 21, 2026, with Hugging Face weights, an official API, and overseas real-time pricing starting at $0.28 per 1 million output tokens for Flash. Xiaomi’s launch announcement is available through its official MiMo V2.6 release post.
Key Takeaways
- MiMo V2.6 is a model family, not one single checkpoint: the main releases are Pro and Flash.
- Both models accept text, images, video, and audio, and support a 1-million-token context window.
- Xiaomi released Pro-RL and Flash-RL weights under the MIT license on Hugging Face.
- Overseas API pricing is $0.435 per 1 million uncached input tokens and $0.87 per 1 million output tokens for Pro.
- Flash costs $0.14 per 1 million uncached input tokens and $0.28 per 1 million output tokens.
- Xiaomi reports a score of 46 on the Artificial Analysis Intelligence Index for Pro, but independent benchmark coverage remains limited.
What Is Xiaomi MiMo V2.6?
MiMo V2.6 is Xiaomi’s latest large-model family, designed around full-modality understanding and complex professional workflows. Its primary models combine language reasoning with visual, video, and audio inputs. The official model listing also includes deep thinking, function calling, structured output, web search, and streaming output.
The family currently has two central open-weight checkpoints:
- MiMo-V2.6-Pro-RL: the flagship model for complex projects, research, cybersecurity, and long-running tasks.
- MiMo-V2.6-Flash-RL: the efficiency-oriented model for frequent calls, professional office work, and large-scale processing.
Xiaomi’s official API documentation lists mimo-v2.6-pro-ultraspeed as a third service variant. Its relationship to the downloadable Pro checkpoint is not fully described in the supplied release material, so it is best treated as an API-serving variant rather than a separate confirmed open-weight release.
The model family is released, not merely leaked or previewed. Xiaomi has published model cards, API documentation, pricing, and a technical report. The release also follows a public reinforcement-learning run that Xiaomi says lasted fewer than 6 days and completed 30 training steps for both models.
Xiaomi MiMo V2.6 at a Glance
| Specification | MiMo-V2.6-Pro | MiMo-V2.6-Flash | MiMo-V2.6-Pro-Ultraspeed |
|---|---|---|---|
| Developer | Xiaomi MiMo | Xiaomi MiMo | Xiaomi MiMo |
| Type | Open-weight native omnimodal model | Open-weight native omnimodal model | API serving variant |
| Modality | Text, image, video, audio understanding; agent and tool use | Text, image, video, audio understanding; agent and tool use | Full-modal understanding listed; additional deployment details not yet confirmed |
| Context window | 1 million tokens | 1 million tokens | 1 million tokens |
| Maximum output | 128,000 tokens | 128,000 tokens | 128,000 tokens |
| Parameter scale | Community-reported 1.02 trillion total / 42 billion active; not independently verified | Community-reported 309 billion total / 15 billion active; not independently verified | Not yet confirmed |
| Overseas real-time pricing | $0.0036 per 1M cache-hit input; $0.435 per 1M cache-miss input; $0.87 per 1M output | $0.0028 per 1M cache-hit input; $0.14 per 1M cache-miss input; $0.28 per 1M output | $0.036 per 1M cache-hit input; $4.35 per 1M cache-miss input; $8.70 per 1M output |
| Availability | Official API and Hugging Face weights | Official API and Hugging Face weights | Official API; weights not yet confirmed |
| License | MIT for the Pro-RL model card | MIT for the Flash-RL model card | Not yet confirmed |
How Xiaomi MiMo V2.6 Works / What Makes It Different
The central design idea is Scaled RL: Xiaomi expanded reinforcement-learning compute, task environments, and grader capacity together. The goal is not simply to improve a static benchmark score. It is to train the model through repeated exploration, tool use, feedback, and task completion.
Xiaomi reports that each RL update used 1,568 samples and that training supported a 1-million-token sequence length. The company also describes a mixed task environment spanning code, general reasoning, visual work, cybersecurity, and chat. The release note reports approximately 3.5 to 3.7 billion training tokens per step and about 750,000 total trajectories across the run.
This training approach produced the Long-Context 1M capability: a model context large enough for lengthy repositories, document collections, extended agent traces, and multi-step project state. A large context window does not guarantee perfect recall across every token, but it changes the size of workloads that can fit into one session.
MiMo V2.6 is also described as a Native Omnimodal model. That means text, image, video, and audio understanding are presented as core model capabilities rather than separate products that must be stitched together. The available documentation establishes multimodal understanding; it does not establish that MiMo natively generates arbitrary images, videos, or audio.
The family adds Computer Use Agent capability, combining visual interface understanding with actions in productivity and software environments. Xiaomi’s published examples include information retrieval, editing, data processing, result checking, and follow-up corrections based on visual feedback.
A related concept in Xiaomi’s release material is Vibe World. In this workflow, natural-language instructions expand from code generation into interactive 3D world creation. Xiaomi describes multi-agent task decomposition for building game scenes, writing interaction logic, rendering the result, and revising the output.
Parameter counts require more care. Early community reports describe Pro as a sparse mixture-of-experts model with approximately 1.02 trillion total parameters and 42 billion active parameters, while Flash is described at 309 billion total and 15 billion active. Those figures are useful for understanding the reported scale, but they are not treated here as independently verified specifications.
Xiaomi also reports engineering controls for RL stability, including router freezing to limit expert-load drift and defenses against reward hacking. The associated technical report and resources are attached to the official MiMo-V2.6-Pro-RL model card.
What You Can Do With Xiaomi MiMo V2.6
MiMo V2.6 is aimed at tasks that combine reasoning, perception, and action.
Coding and software agents: Pro is positioned for complex software projects, while Flash is intended for higher-volume calls. The 1-million-token context can hold larger codebases or longer execution histories than ordinary short-context chat models.
Computer-use workflows: The model can interpret interfaces and operate common productivity tools according to Xiaomi’s examples. This includes retrieving information, editing documents, processing data, checking visual results, and adjusting later actions.
3D creation: Xiaomi demonstrates text- and image-guided Blender scene and object creation. The output can support animation, 3D printing, and game development workflows, although production quality still requires human review.
Interactive world building: The Vibe World workflow connects planning, code generation, scene construction, rendering, and visual verification. It is a practical fit for prototypes that need both software logic and visual assets.
Scientific and technical research: Xiaomi reports examples involving materials research, literature and patent review, computational screening, and Lean 4 formalization. One reported formalization project produced more than 6,000 lines of Lean code and passed kernel verification after researcher revision and integration.
These examples describe documented demonstrations and workflows, not guarantees for every prompt. The most reliable use cases will still need sandboxing, tool permissions, evaluation sets, and human review.
How Xiaomi MiMo V2.6 Compares
Xiaomi positions MiMo V2.6 Pro as an open-weight alternative to frontier closed models. Its official release reports a score of 46 on the Artificial Analysis Intelligence Index and says Pro exceeds Kimi K3 and Qwen3.8 Max in that comparison. Xiaomi also says the model remains behind the strongest closed models, including Claude Fable 5.1 and GPT-6 Astra.
| Comparison point | MiMo V2.6 signal | Confidence |
|---|---|---|
| Artificial Analysis Intelligence Index | Pro scored 46, according to Xiaomi | Official claim; independent methodology review remains limited |
| Open-model comparison | Xiaomi says Pro exceeds Kimi K3 and Qwen3.8 Max | Vendor-reported comparison |
| Closed-model comparison | Xiaomi says Pro still trails Claude Fable 5.1 and GPT-6 Astra | Vendor-reported comparison |
| Open deployment | Pro-RL and Flash-RL weights are available under MIT model cards | Confirmed for the listed checkpoints |
For broader context, the trade-off between open weights and frontier hosted performance is also discussed in the earlier analysis of Claude Fable 5.1’s context and pricing and the GPT-6 Astra overview.
The defensible verdict is narrower than “best model.” MiMo V2.6 is unusually broad for an open-weight release, combining 1-million-token context, four input modalities, agent tooling, and low listed API prices. Its overall ranking still needs more reproducible third-party testing.
Availability: How to Access Xiaomi MiMo V2.6
There are two main access paths.
Official API: Xiaomi’s MiMo API platform supports OpenAI-compatible and Anthropic-compatible request formats. The official model list names mimo-v2.6-pro, mimo-v2.6-flash, and mimo-v2.6-pro-ultraspeed. Pay-as-you-go billing uses ordinary API credentials, while Xiaomi also lists a separate Token Plan subscription.
The listed overseas real-time rates are $0.87 per 1 million output tokens for Pro and $0.28 for Flash. Batch API pricing is listed at 50% of real-time rates for Pro and Flash. Web search is billed separately from token usage.
Downloadable weights: Xiaomi publishes Pro-RL and Flash-RL model cards on Hugging Face. The cards identify the MIT license and provide deployment paths using Transformers, vLLM, SGLang, Docker, and compatible local applications. The Flash model card is available at XiaomiMiMo/MiMo-V2.6-Flash-RL.
Local deployment is not lightweight. Community reports place the shipped FP8 files at approximately 177.8 gigabytes for Flash and 573.5 gigabytes for Pro, but hardware requirements, throughput, and quantization behavior depend on the serving stack. Builders evaluating hosted alternatives for comparable reasoning workflows can also try Claude Opus 5.
What We Don't Know Yet
Several practical questions remain open:
- The exact architecture and parameter accounting for every variant are not fully consolidated in the supplied public material.
- The relationship between Pro-Ultraspeed and the downloadable Pro-RL checkpoint is not clearly documented.
- Independent evaluations with complete prompts, costs, confidence intervals, and contamination checks are still limited.
- Minimum hardware, multi-GPU layouts, throughput, and long-context memory usage need reproducible testing.
- The documentation confirms multimodal understanding, but it does not establish native image, video, or audio generation.
- Global API coverage, regional restrictions, quotas, and Token Plan terms may change as the service matures.
Frequently Asked Questions
What is Xiaomi MiMo V2.6?
Xiaomi MiMo V2.6 is Xiaomi’s open-weight family of native omnimodal AI models for text, image, video, audio, reasoning, coding, and agent workflows. The primary Pro and Flash models support a 1-million-token context window and are available through official APIs and downloadable model weights.
Is Xiaomi MiMo V2.6 open source?
Xiaomi MiMo V2.6 is available as open weights under the MIT license for the Pro-RL and Flash-RL checkpoints hosted on Hugging Face. That permits commercial use and further training under the stated license, although deployment requirements remain substantial.
How much does Xiaomi MiMo V2.6 cost?
Xiaomi MiMo V2.6 costs $0.87 per 1 million output tokens for Pro and $0.28 per 1 million output tokens for Flash on the listed overseas real-time API. Input pricing is $0.435 per 1 million uncached tokens for Pro and $0.14 for Flash, with lower cache-hit rates.
What is the context window of Xiaomi MiMo V2.6?
Xiaomi MiMo V2.6 has a 1-million-token context window, with a maximum output length of 128,000 tokens in the official model listing. The same limit is listed for Pro, Flash, and Pro Ultraspeed API variants.
How can I access Xiaomi MiMo V2.6?
Xiaomi MiMo V2.6 can be accessed through Xiaomi’s official MiMo API platform or by downloading the Pro-RL and Flash-RL weights from Hugging Face. The model cards document Transformers, vLLM, SGLang, and Docker-based deployment paths.
Xiaomi MiMo V2.6 vs Claude Opus 5?
Xiaomi MiMo V2.6 is the more accessible option for builders who need open weights, multimodal inputs, and MIT-licensed deployment. Xiaomi reports performance near Claude Opus 5 on many agent benchmarks, but independent testing has not yet established a definitive overall winner.
What to watch next is independent testing of the 1-million-token workflows, a clearer release specification for Pro-Ultraspeed, and reproducible reports on local inference cost and throughput. Those signals will determine whether MiMo V2.6’s open-weight advantage translates into dependable production use.
Building similar long-context multimodal reasoning and agent workflows? On kie.ai you can try Claude Opus 5.5, Kimi K3, and GPT 6.1 Sol.
About Daniel Okonkwo
Daniel writes about inference systems, model architecture, and what new releases actually change for builders.
View all posts by Daniel Okonkwo