Mistral Large 4 101: The 1M-Token, 1.05T-Parameter MoE Model

Elena Rossi

Elena Rossi

AI Adoption Analyst

Published: October 6, 2026
Mistral Large 4 Le Chonk multimodal mixture-of-experts model

TLDRA 1.05T-parameter, 49B-active multimodal MoE with a 1M-token context window, API preview access, and planned open weights.

Mistral Large 4 is Mistral AI’s open-weight, general-purpose multimodal mixture-of-experts model with 1.05 trillion total parameters, 49 billion active parameters, and a 1 million-token context window. Mistral released it as a public preview on October 6, 2026, with API access through Mistral Studio and downloadable weights planned for the end of October. Its internal nickname is “Le Chonk,” and its initial positioning centers on coding, AI agents, cybersecurity, finance, manufacturing, and visual understanding.

Key Takeaways

  • Mistral Large 4 is a native text-and-image model built by Mistral AI.
  • Its architecture contains 1.05 trillion total parameters, with 49 billion active for each token.
  • The model supports a 1 million-token context window and includes a 1.6 billion-parameter vision encoder.
  • The public preview is available through Mistral Studio at listed prices of $0.68 per 1 million input tokens and $2.09 per 1 million output tokens.
  • Mistral says the model reaches state-of-the-art results among open models in cybersecurity, finance, and manufacturing.
  • The API is available now, but the weights, license, and practical self-hosting requirements are not fully published.

What Is Mistral Large 4?

Mistral Large 4 is Mistral AI’s flagship open-weight model for language, images, reasoning, code, and agentic workflows. Mistral describes it as a hybrid instruct-and-reasoning MoE model, meaning it combines instruction following and deliberate problem solving in a sparse mixture-of-experts architecture.

The model entered public preview on October 6, 2026. Mistral’s official announcement says that the preview API is available through Mistral Studio and that the weights will drop by the end of the month. That makes Large 4 released as an API preview, but not yet fully released as a downloadable checkpoint.

Mistral developed and trained the model from scratch in its European data centers. The company says training used 3,800 NVIDIA Grace Blackwell GPUs. The public preview is served on the same European infrastructure, and Mistral says the model can be deployed in Europe under European law.

Mistral Large 4 is therefore aimed at two audiences. API users can test a frontier-scale model without operating its infrastructure. Enterprises and researchers waiting for the weights can eventually inspect, customize, and self-deploy the model, subject to the final license and hardware requirements.

Mistral Large 4 is an API-preview model today and a planned open-weight checkpoint for the end of October 2026.

Mistral Large 4 at a Glance

SpecificationMistral Large 4
DeveloperMistral AI
TypeOpen-weight hybrid instruct-and-reasoning MoE
Total parameters1.05 trillion
Active parameters49 billion per token
ModalityText and image input
Vision encoder1.6 billion parameters
Context window1 million tokens
LanguagesMore than 160 languages
API price$0.68 per 1 million input tokens; $0.07 cached input; $2.09 output
AvailabilityPublic preview through Mistral Studio
Downloadable weightsPlanned for the end of October 2026
LicenseNot yet confirmed
Self-hosting requirementsNot yet confirmed

The pricing figures above come from Mistral’s current pricing documentation. Mistral’s model page also displays a higher $1.36 input price, $0.14 cached-input price, and $4.18 output price for another listed inference tier. Developers should check the live pricing page before estimating production costs.

How Mistral Large 4 Works and What Makes It Different

The central technical choice is sparse activation. Mistral Large 4 has more than one trillion parameters in total, but only 49 billion parameters are active for each token. This MoE design gives the model a large pool of learned capabilities without requiring every parameter to run on every token.

Mistral Large 4 is a sparse MoE: 1.05 trillion total parameters, but 49 billion active for each token.

That distinction matters for engineering. Total parameter count describes the model’s overall capacity and storage footprint. Active parameter count is more relevant to the computation used during each inference step. It does not make the model small or easy to run locally, because the full expert set still has to be stored and managed.

The model’s Native Multimodality supports text and image input in one general-purpose model. Mistral also lists a 1.6 billion-parameter vision encoder. This design targets document question answering, screenshot analysis, visual grounding, and workflows where an agent must combine written instructions with images.

The model’s Hybrid Instruct-and-Reasoning identity is another important anchor. It is not presented as a separate reasoning-only checkpoint. Instead, Mistral positions instruction following, reasoning, coding, and agent behavior inside a single model.

Its 1 million-token context window supports long repositories, large document collections, extended agent traces, and multi-file engineering tasks. Context length is a capacity limit, not a guarantee of perfect recall. Quality still depends on prompt structure, retrieval strategy, token budget, and the model’s ability to locate relevant information.

Mistral also emphasizes European Sovereign Deployment. The model was trained on Mistral’s own infrastructure in Europe, and the preview is served there. This positioning is especially relevant to organizations with data residency, regulatory, or operational-control requirements.

Finally, Mistral calls out Visual Grounding, cybersecurity, finance, manufacturing, and legal work. The company claims that Large 4 surpasses closed frontier models in some visual-grounding tests. That is a vendor claim awaiting broader independent verification.

What You Can Do With Mistral Large 4

Mistral Large 4 is designed for workloads that need a broad model rather than a narrow specialist. The strongest early use cases are:

  • Software engineering: code generation, repository analysis, debugging, terminal-style tasks, and software agents.
  • Cybersecurity: vulnerability research, incident response, defensive analysis, and security operations under organization-specific policies.
  • Multimodal document work: extracting meaning from diagrams, screenshots, scanned documents, and image-heavy technical material.
  • Enterprise agents: planning, tool use, structured outputs, function calling, and long-running conversations.
  • Finance and legal analysis: research, document review, financial workflows, and domain-specific agents.
  • Manufacturing and engineering: interpreting technical documents, images, procedures, and operational records.
  • Multilingual applications: Mistral says the training data covered more than 160 languages, including every official language of the European Union.

The official model documentation lists structured outputs, function calling, document question answering, batching, agents, conversations, and built-in tools. These features make the model relevant to production application design, although preview status means interface behavior and limits may still change.

A 1 million-token context window can reduce the need to aggressively summarize a large repository or document set. It does not remove the need for retrieval, chunking, access controls, or evaluation. Teams should test long-context accuracy against their own data before assuming that the maximum window improves every workflow.

For a separate comparison point while testing agent-oriented chat workflows, GPT-6 Astra is another cataloged model aimed at demanding general-purpose tasks.

How Mistral Large 4 Compares

Mistral Large 4 is being compared most often with GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash. The early numbers are useful directionally, but they do not form one standardized leaderboard.

ModelEarly comparison signalEvidence status
Mistral Large 461.7% on DeepSWE v1.1; 50 on the Artificial Analysis Cyber IndexMistral-reported or early community-reported
GLM 5.361% on a cited DeepSWE comparison; 36 on the cited Cyber IndexMixed reported sources
Kimi K3About 68% on the cited DeepSWE comparisonEarly community report
DeepSeek V4.1 FlashArtificial Analysis Intelligence Index score of 39 in the cited comparisonThird-party report

Early community testing places Mistral Large 4 around GLM 5.3 on coding, while Kimi K3 is reported higher on one DeepSWE comparison. Another early report gave Mistral Large 4 an Artificial Analysis Intelligence Index score of 38, one point below DeepSeek V4.1 Flash at 39.

Early benchmark evidence places Mistral Large 4 near GLM 5.3 on coding, not conclusively above every frontier model.

The comparisons have important limits. Some scores come from Mistral’s own evaluation tables, while others come from third-party trackers or community screenshots. Test harnesses, prompts, inference settings, tool use, and output-token budgets may differ. The scores should guide testing priorities, not replace application-specific evaluation.

Mistral’s more defensible distinction may be the combination of capabilities, European deployment, and planned weights. A model can be commercially useful without ranking first on every public benchmark, particularly when customers prioritize control, data residency, or customization.

Availability: How to Access Mistral Large 4

Mistral Large 4 is available as a public preview through Mistral Studio. The official model documentation identifies the model as PUBLIC PREVIEW and lists chat completions, conversations, agents, function calling, structured outputs, document question answering, and batching.

The listed API pricing is:

  • $0.68 per 1 million input tokens
  • $0.07 per 1 million cached input tokens
  • $2.09 per 1 million output tokens

Mistral also displays a second pricing tier at $1.36 per 1 million input tokens, $0.14 per 1 million cached input tokens, and $4.18 per 1 million output tokens. The exact applicable tier depends on the inference configuration shown in Mistral’s documentation.

Downloadable weights are not available in the launch-day preview. Mistral’s announcement says “by the end of the month.” The Hugging Face model page shows an expected October 31, 2026 release, while several secondary reports cite October 27. The final date should be treated as unconfirmed until the checkpoint is published.

What We Don't Know Yet

Several practical details remain open:

  • The final open-weight license has not been confirmed.
  • The downloadable checkpoint, quantization options, and supported inference software are not yet published.
  • Minimum hardware requirements for self-hosting are not confirmed.
  • The exact model architecture beyond the MoE description and vision encoder has not been fully documented.
  • The official announcement does not specify whether the 1 million-token context window applies identically across every access tier.
  • Independent, reproducible benchmark results remain limited.
  • The reported 3,800-GPU training figure differs from secondary accounts that round the number to 4,000 GPUs.
  • Production service-level guarantees, rate limits, and regional API details are not established in the available material.
  • Final weight-release timing remains a schedule rather than a completed release.

Mistral says it is red-teaming the model with cybersecurity leaders, vetted partners, and state authorities before releasing the weights. That process may affect moderation, cyber capabilities, and the final downloadable configuration.

Frequently Asked Questions

What is Mistral Large 4?

Mistral Large 4 is Mistral AI’s open-weight, general-purpose multimodal mixture-of-experts model with 1.05 trillion total parameters, 49 billion active parameters, and a 1 million-token context window. It entered public preview through Mistral Studio on October 6, 2026.

Is Mistral Large 4 open source?

Mistral Large 4 is an open-weight model, but its downloadable weights and final license terms are scheduled for release at the end of October 2026 and are not yet fully confirmed. “Open-weight” does not automatically describe the final permissions for commercial use, redistribution, or modification.

How much does Mistral Large 4 cost?

Mistral’s pricing page lists Mistral Large 4 at $0.68 per 1 million input tokens, $0.07 per 1 million cached input tokens, and $2.09 per 1 million output tokens for the listed inference tier. The same documentation displays higher prices for another tier, so production estimates should use the applicable current rate.

What is the context window of Mistral Large 4?

Mistral Large 4 has a 1 million-token context window according to Mistral’s model documentation. The practical quality of long-context tasks still depends on prompt design, retrieval, and the model’s ability to find relevant information.

How can I access Mistral Large 4?

Developers can access Mistral Large 4 through the public preview API in Mistral Studio. Downloadable weights are planned for the end of October 2026, but the exact release date and final license are not yet confirmed.

Mistral Large 4 vs GLM 5.3: which is better?

Mistral Large 4 is near GLM 5.3 on early coding tests, but the available comparisons use different evaluation sources and do not establish a universal winner. Mistral Large 4 has a reported DeepSWE v1.1 result of 61.7%, while one cited GLM 5.3 result is 61%.

What to Watch Next

The most important updates are the actual weight release, the final license, and independent tests of long-context, coding, cyber, and visual-grounding performance. Developers should also watch whether Mistral publishes full architecture details, self-hosting guidance, and stable production pricing.

Building similar multimodal agent workflows? On kie.ai you can try Claude Sonnet 5.5, Kimi K3, and DeepSeek-V4.1-Flash.

Elena Rossi

About Elena Rossi

Elena watches developer chatter and early adoption signals to gauge which releases gain real traction.

View all posts by Elena Rossi