What Is Kimi K3? Moonshot's 2.8T, 1M-Context Flagship

Daniel Okonkwo

Daniel Okonkwo

Senior ML Engineer

Published: July 15, 2026
Leaked Moonshot AI promo page announcing the Kimi K3 launch on July 15, 2026

TLDRKimi K3 is Moonshot AI's next-gen model — a ~2.8T MoE with a 1M-token context window, launched July 16, 2026 on Kimi Code and the Kimi app.

What Is Kimi K3? Moonshot’s 2.5T-Parameter, 1M-Context Flagship

Kimi K3 is Moonshot AI's next-generation large language model, launched on July 16, 2026 after a leaked promotion page on Moonshot's own Kimi Open Platform tipped the release a day early. It is a new-architecture Mixture-of-Experts model with roughly 2.8 trillion total parameters and a 1-million-token context window, aimed at long-horizon coding and agent workloads. Two variants shipped at launch — K3 Max for chat and agent tasks, and K3 Swarm Max for large-scale parallel processing — first on Kimi Code and inside the Kimi app, with an accompanying top-up promotion offering 10–30% bonus credits on API recharges through August 11.

Updated 2026-07-16: Kimi K3 has officially launched at ~2.8T parameters with 1M context, rolling out on Kimi Code and the Kimi app (see the Update below).

Key Takeaways

  • Kimi K3 is Moonshot AI's flagship model, launched July 16, 2026 as the successor to the Kimi K2 family (K2, K2.5, K2.6, K2.7 Code).
  • Confirmed at launch: ~2.8T total parameters, 1M-token context window, open-source release, two variants (K3 Max and K3 Swarm Max).
  • Available now on Kimi Code and inside the Kimi app (including iOS), with a ¥199 subscription tier as the entry point.
  • The launch is accompanied by a recharge promotion running July 15 through August 11 with 10–30% bonus credits.
  • Moonshot's $500M Series C, closed January 2026 at a $4.3B valuation, was explicitly earmarked for K3 development and compute expansion.
  • Early users describe K3 as competitive with top-tier closed coding models, with independent estimates around Opus 4.8 / GPT-5.5 tier on Artificial Analysis and behind Fable 5 on some arena prompts.

What Is Kimi K3?

Kimi K3 is the third-generation model in Moonshot AI's Kimi series, the successor to the Kimi K2 family that shipped between July 2025 and mid-2026. Moonshot AI is a Beijing-based startup best known for the Kimi chatbot and for releasing K2 under a Modified MIT license, which made it one of the strongest open-weight models of 2025.

CEO Yang Zhilin publicly acknowledged K3 as the company's next major generation earlier in 2026. Moonshot's January 2026 $500 million Series C at a $4.3 billion valuation was reported to fund "computing capacity and developing the K3 model," per The Decoder's coverage of the round.

Kimi K3 went live on July 16, 2026. The model rolled out first on Kimi Code and inside the Kimi app, with two variants surfaced at launch — K3 Max for chat and agent tasks, and K3 Swarm Max for large-scale parallel processing. A leaked promotion page on Moonshot's own API platform had tipped a July 15 (China time) launch a day earlier; the actual rollout began on the evening of July 16.

🚨 Kimi K3 set to launch within hours Moonshot AI may have accidentally revealed the date through it

Source: @LuminaXspace

Kimi K3 at a Glance

AttributeValue
DeveloperMoonshot AI
TypeLarge language model, Mixture-of-Experts
ModalityText (multimodal scope not yet detailed)
Total parameters~2.8T
Active parametersNot yet disclosed
Context window1M tokens
ArchitectureNew architecture, distinct from K2 series
LicenseOpen source (weights expected in the coming days; K2 family used Modified MIT)
API pricingPer-token rates not yet published; ¥199 subscription tier available
Launch dateJuly 16, 2026
AvailabilityKimi Code, Kimi app (including iOS); two variants: K3 Max and K3 Swarm Max
PredecessorKimi K2.7 Code / K2.6 (1T MoE, 32B active, 256K context)

Cells marked "not yet disclosed" or "not yet published" reflect gaps Moonshot hasn't ratified at launch.

How Kimi K3 Works and What Makes It Different

Kimi K3 is described as built on a "new architectural innovation" rather than a straight scale-up of the K2-series MoE. The Kimi K2 family, by contrast, keeps a stable blueprint across versions: 1 trillion total parameters, 32 billion active per token, 384 routed experts plus 1 shared expert, Multi-head Latent Attention, and SwiGLU activation with native INT4 quantization, per a third-party K2 architecture writeup that summarizes Moonshot's published technical reports.

For K3, several distinct architectural hints appear across pre-launch reporting:

  • Community researcher Teortaxes speculated on hardware and algorithmic choices, saying "I expect it to be the most advanced base model yet. I want to see AttnRes, modality vision, Kimi-Linear+ on 2T+ scale," in a post on July 7, 2026.
  • One commenter on the r/kimi subreddit noted Moonshot has "been publishing papers that mentioned '1T internal hybrid linear model,'" suggesting exploration of hybrid linear attention variants, in the Reddit launch thread.
  • Leaked spec-sheet screenshots circulating on July 14 referenced terms like DSA and "Kimi residual attention," with claims of the sparsest MoE ratio yet and native 4-bit training. Moonshot has not yet published a full model card confirming these details.

The consistent thread across the launch positioning is a focus on long-horizon agent tasks, large-scale parallel search and execution (materialized in the K3 Swarm Max variant), and stronger coding stamina. This lines up with Moonshot's direction in K2.6 and K2.7 Code, which already emphasized multi-hour autonomous runs and thousands-of-tool-calls sessions.

Three anchor terms recur enough across the release to be worth tracking as the model's identity crystallizes: K3 Swarm Max, long-horizon agent workloads, and new architectural innovation. Only the first now points to a concrete shipping product.

What You Can Do With Kimi K3

Because K3 just launched, use cases are anchored in Moonshot's positioning, K2.7 Code's actual capabilities, and early-user reports on July 16.

  • Long-session autonomous coding. K2.6 already demonstrated 12–13 hour autonomous coding runs with thousands of tool calls, and K3 extends this direction under the K3 Swarm Max variant for large-scale parallel processing.
  • Large-context research and analysis. With the confirmed 1M-token context window, K3 sits in the same tier as long-context flagships for whole-repository code review, long document synthesis, and multi-document research.
  • Coding agent backends. Early tester ChrissGPT positioned K3 pre-launch as "an Opus 4.7+ coding model" that "outperforms GPT 5.5 and GPT 5.6 Terra on SOME coding Evals," while noting that Fable 5 and GPT-5.6 Sol still dominate on Terminal-Bench 2.1, in a July 14 post. At launch, he pegged K3 at Opus 4.8 level in coding.
  • Multi-agent orchestration. K2.6 shipped Claw Groups for coordinating heterogeneous external agents; K3 Swarm Max continues this direction with explicit large-scale parallel processing support.

Treat these as directional until independent benchmarks land and Moonshot publishes a full model card.

How Kimi K3 Compares

The most useful comparison is against Kimi's own K2 family, since that is the only apples-to-apples baseline available.

ModelTotal paramsActive paramsContextLicenseStatus
Kimi K2 (Jul 2025)1T32B256KModified MITReleased, EOL May 2026
Kimi K2.6 (Apr 2026)1T32B256KModified MITReleased
Kimi K2.7 Code (Jun 2026)1T32B256KModified MITReleased
Kimi K3 (Jul 2026)~2.8TNot yet disclosed1MOpen source (details pending)Released

K3 is also being read against contemporary releases — DeepSeek V4 and the rumored GLM-5.x are the most-cited comparisons in community threads. For the V4 context, see our earlier analysis of the DeepSeek V4 release signals. For the parallel flagship-frontier picture at Western labs, our writeup on the Grok 4.5 leak and the 1.5T V9 foundation covers the same competitive window.

Availability: How to Access Kimi K3

Kimi Code: Live. Ivan Fioravanti confirmed the model was available on Kimi Code at rollout, in a July 16 post.

Kimi app (web and iOS): Live. Mark Kretschmann confirmed the model was accessible via Kimi AI at launch, and K3 Max is available on the iOS app.

Two variants at launch: K3 Max for chat and agent tasks, and K3 Swarm Max for large-scale parallel processing, per @LuminaXspace's rollout note and @ivanfioravanti's screenshots.

Open source: Yes. K3 launched as an open-source model. Full weights are expected to be released in the coming days per early-launch reports; K2's precedent is Modified MIT.

Public API pricing: Per-token rates are not yet published. LinearUncle noted a ¥199 subscription tier as an entry point, in a July 16 post. The launch is accompanied by a "Kimi K3 launch limited-time recharge campaign" running July 15 through August 11, with Chubby posting the promo structure: "¥99–¥499: 10% bonus / ¥500–¥1,999: 20% bonus / ¥2,000–¥4,999: 25% bonus / ¥5,000 or more: 30% bonus," in a July 14 post.

For local inference, developer Max Weinbach flagged the hardware ceiling early: "I'm getting my mac studio cluster ready for Kimi K3 but I'm pretty sure I'm hitting the limits of what 1.5TB if memory can do for a 2.5T parameter 1M context model," in a July 14 post. At the confirmed ~2.8T total, K3 stretches even top-end workstation setups, and third-party inference throughput remains an open question.

What We Don't Know Yet

Even with launch, several gaps remain as of July 16, 2026:

  • Sparsity ratio and active parameter count. Moonshot has not disclosed how many parameters activate per token, which drives real inference cost.
  • Full architecture details. Whether "new architecture" means hybrid linear attention, a different expert routing scheme, DSA-style variants, or something else — the leaked terms have not been officially confirmed.
  • Multimodal scope. Pre-launch rumors mentioned image, audio, and text input; the launch surface is text-first, and full multimodal details are pending.
  • Per-token API pricing. Only the recharge promotion structure and the ¥199 subscription tier are visible; per-MTok input/output rates are unpublished.
  • Independent benchmarks. No third-party evaluations on SWE-Bench Verified, Terminal-Bench 2.1, HLE, or agent-specific harnesses yet.
  • Weights availability. Open-source is confirmed; the actual Hugging Face repo, license text, and configuration file are expected in the coming days.
  • Inference-provider throughput. As Max Weinbach flagged post-launch, whether providers can scale a ~2.8T model to competitive tokens-per-second economics is the practical unknown.

Frequently Asked Questions

What is Kimi K3?

Kimi K3 is the next-generation large language model from Moonshot AI, the Beijing-based company behind the Kimi chatbot and the Kimi K2 open-weight family. Moonshot began rolling out K3 on July 16, 2026, first on Kimi Code and inside the Kimi app, with two variants at launch: K3 Max for chat and agent tasks, and K3 Swarm Max for large-scale parallel processing.

When was Kimi K3 released?

Kimi K3 launched on July 16, 2026, when Moonshot AI began rolling it out on Kimi Code and inside the Kimi app. The launch followed a leaked July 15 (China time) promotion page that overshot by roughly a day, and multiple early users confirmed live access on July 16.

How many parameters does Kimi K3 have?

Kimi K3 has approximately 2.8 trillion total parameters, above the 2.5T figure that dominated pre-launch leaks. It is a Mixture-of-Experts model, though Moonshot has not yet disclosed the sparsity ratio or active parameter count.

Is Kimi K3 open source?

Yes. Kimi K3 launched as an open-source model, continuing the open-weight pattern that Moonshot AI established with the Kimi K2 family under a Modified MIT license. Full weights are expected to be released in the coming days per early-launch reports.

How much does Kimi K3 cost?

Per-token API pricing for Kimi K3 has not been published. The launch is accompanied by a limited-time recharge promotion offering 10–30% bonus credits on top-ups between ¥99 and ¥5,000+, running July 15 through August 11, and a ¥199 subscription tier is available on the Kimi platform. For reference, Kimi K2.6 launched at $0.60 input / $2.50 output per million tokens.

What is Kimi K3's context window?

Kimi K3 supports a 1-million-token context window, confirmed at launch by multiple early users and matching the pre-release leaks.

Kimi K3 vs Kimi K2 — what's different?

The Kimi K2 family is a 1-trillion-parameter MoE with 32B active parameters and a 256K context window, released under a Modified MIT license in July 2025. Kimi K3, launched July 16, 2026, jumps to roughly 2.8T total parameters and a 1M-token context window, on a new architecture targeting long-horizon agent workloads.

Where can I try Kimi K3?

Kimi K3 is live on Kimi Code and inside the Kimi app, including the Kimi iOS app, as of July 16, 2026. Two variants are available: K3 Max for chat and agent tasks, and K3 Swarm Max for large-scale parallel processing.

What to Watch Next

Three concrete signals will move this page from "launched" to "fully specified": a Moonshot AI model card that ratifies the sparsity ratio, active parameter count, and full architecture; a Hugging Face repository with weights and license text; and independent benchmark runs on SWE-Bench Verified and Terminal-Bench 2.1 that test the "Opus 4.8-level coding" claim against a comparable harness. This page will be updated in place as each lands.

Update — 2026-07-16

Moonshot AI began rolling out Kimi K3 on the evening of July 16, 2026, ending the leak cycle covered above. The model went live first on Kimi Code and inside the Kimi app, with two variants surfaced at launch: K3 Max for chat

Daniel Okonkwo

About Daniel Okonkwo

Daniel writes about inference systems, model architecture, and what new releases actually change for builders.

View all posts by Daniel Okonkwo