Kimi K3 Release: Signal vs Noise on Moonshot's 2.8T Model
Elena Rossi
AI Adoption Analyst

TLDRMoonshot's Kimi K3 shipped as a 2.8T open-weight model with 1M context. What the signals confirm, what's still unverified.
Kimi K3 Release: Signal vs Noise on Moonshot's 2.8T Open Frontier Model
Moonshot AI released Kimi K3 on July 16, 2026, then shipped its full weights to Hugging Face eleven days later on July 27 — a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window and native vision, positioned as the largest open-weight model ever published.
TLDR Kimi K3 is a 2.8T-parameter MoE built on a new attention stack called Kimi Delta Attention (KDA) plus Attention Residuals, activating 16 of 896 experts per token. It launched hosted on Moonshot's platforms with a 1M-token context, then opened its weights on Hugging Face on July 27, priced at $3/$15 per million tokens on the official API. Independent benchmarks from Artificial Analysis rank it fourth on the Intelligence Index — behind Claude Opus 5, Claude Fable 5, and GPT-5.6 Sol, but ahead of Grok 4.5 and every other open-weight model. A UK AISI / CAISI joint assessment reports it lags US frontier models on cyber-offense benchmarks.
Key Takeaways
- Kimi K3 is a 2.8T-parameter MoE with 16 of 896 experts active per token and ~2.5× better scaling efficiency than Kimi K2, per Moonshot's official blog.
- New architecture components: Kimi Delta Attention (KDA), Attention Residuals (AttnRes), and a Stable LatentMoE framework.
- Open weights shipped July 27, 2026 to Hugging Face — exactly on the date Moonshot committed to at launch.
- API pricing is $3 input / $15 output per million tokens on Moonshot's platform, with a 90% cache-hit discount to $0.30.
- Independent Artificial Analysis Intelligence Index score of 57, ranking fourth overall and first among open-weight models in its size class.
- Day-0 hosting partners include Nebius, Baseten, Fireworks AI, DigitalOcean, and Together AI.
- Not confirmed: a full technical report, license details for commercial deployment, and third-party reproductions of Moonshot's kernel-optimization and compiler-development case studies.
What Actually Shipped on July 16
The rollout began on the evening of July 16, 2026 China time. Ivan Fioravanti caught the Kimi CLI updating twice in short succession — v0.25.0 to v0.26.0 within about eleven hours — and confirmed K3 was live on Kimi Code minutes later. Two variants appeared in the product UI: K3 Max for chat and agent tasks, and K3 Swarm Max for large-scale parallel processing, both surfaced by Lumina in the first hour.
By evening China time, Kimi K3 was available in the Arena.ai Agent Arena for long-horizon agentic tasks alongside Text, Vision, Document, and Frontend Code arenas. Riley Brown reported K3 Max shipping in the iOS app the same afternoon.
The primary source of record is the official Kimi K3 tech blog on kimi.com, which describes the model as "the world's first open 3T-class model" (Moonshot rounds 2.8T up to 3T) built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with the full weights promised for July 27, 2026. That deadline held.
The Architecture in One Paragraph
Per Moonshot's blog and the Hugging Face model card, Kimi K3 pairs two novel attention components — Kimi Delta Attention (a hybrid linear attention mechanism) and Attention Residuals (a mechanism for information flow across model depth) — with a Stable LatentMoE routing framework that activates 16 of 896 experts per token. Moonshot claims a roughly 2.5× improvement in scaling efficiency over Kimi K2, meaning more capability per unit of training compute. Active parameter count is 104B per token according to Artificial Analysis's model page. The model natively accepts vision input and outputs text only.
That's the whole confirmed architecture story. A full technical report is promised but not yet published; anything beyond these bullets — layer counts, exact expert dimensions, RoPE variants, training token budget, data mix — is not in the current signal set.
Pricing and Availability
Moonshot's Kimi API Platform documentation exposes K3 through the OpenAI-compatible /v1/chat/completions endpoint under the model ID kimi-k3. Reasoning effort is controlled via the top-level reasoning_effort field; at launch only max was available. By July 17, TestingCatalog reported the Kimi web and mobile apps had added Standard, High, and Max reasoning effort levels, and LinearUncle noted High became API-configurable a few days later.
Published API pricing, per Artificial Analysis, is $3.00 per 1M input tokens and $15.00 per 1M output tokens, with a cache-hit price of $0.30 — a 90% discount. That places K3 in the same price bracket as Claude Sonnet, not the sub-dollar tier Chinese labs are often assumed to occupy. A Reddit thread on r/OpenAI captured the reaction: "Kimi K3 is $3/$15 per million tokens. That's not cheap Chinese AI anymore".
Day-0 hosting partners announced by Moonshot on the weights-release day included Nebius, Baseten, Fireworks AI, DigitalOcean Serverless Inference, and Together AI. Teortaxes also observed K3 live on Nebius eu-north1 in FP4 at the same cost as full precision but three to four times faster — one of the first quantized deployment signals.
Demand outpaced Moonshot's own capacity almost immediately. TestingCatalog reported Moonshot pausing new subscriptions on July 19 while adding compute — three days after launch, and eight days before the open weights shipped.
Benchmarks: What Independent Numbers Exist
The most credible independent scorecard comes from Artificial Analysis. Kimi K3 lands at 57 on the Artificial Analysis Intelligence Index, ranking fourth overall behind Claude Opus 5 (61), Claude Fable 5 (60), and GPT-5.6 Sol (59). It sits ahead of Grok 4.5 (54), GLM-5.2 (51), Muse Spark 1.1 (51), Gemini 3.6 Flash (50), MiniMax-M3 (44), DeepSeek V4 Pro (44), and Nemotron 3 Ultra (38). Among open-weight models it is currently the highest-scoring entry.
Artificial Analysis flags two nuances. First, K3 is verbose: it generated 130M output tokens during Intelligence Index evaluation, well above the 99M median, which pushes its cost per task to $0.94 — roughly on par with GPT-5.6 Sol's $1.04, despite Kimi's lower per-token sticker. Second, K3 is slow: 33 output tokens per second, near the bottom of the current frontier cohort where GLM-5.2 clears 220 t/s.
Moonshot's own Kimi K3 blog highlights case studies in GPU kernel optimization (across NVIDIA H200 and a second GPGPU vendor) and GPU compiler development (a Triton-like compiler called MiniTriton). Moonshot claims K3 performed competitively with Claude Fable 5 and substantially outperformed Opus 4.8, GPT-5.6 Sol, and GPT-5.5 on the kernel task. These are self-reported and have not yet been reproduced by third parties in the public signal.
One important preliminary independent evaluation came from UK AISI and CAISI's joint cyber assessment, published July 23. They report that Kimi K3 performs significantly below the most recent frontier cyber-capable models on ExploitBench, and on a simulated corporate-network attack path reached step 17 of 32 on average, versus 28.5 for the most cyber-capable US models. K3 outscored GLM-5.2 on the same battery. This is one of the first frontier-lab evaluations from a national safety institute on a Chinese open-weight release.
Kimi K3 vs Claude Fable 5 and GPT-5.6 Sol: What the Signal Says
Two competitors are named repeatedly across the bundle: Claude Fable 5 and GPT-5.6 Sol. Both are treated by community benchmarks as the current proprietary frontier. Kimi K3 is measured against them explicitly by Artificial Analysis, Moonshot's blog, and a widely-shared r/LLMDevs post arguing K3 sits "between the previous frontier tier (Opus 4.8 / GPT-5.5) and the current one (Fable 5 / GPT-5.6 Sol)".
Comparison across five dimensions, sticking to numbers present in the signal set:
- Parameters: Kimi K3 is 2.8T total / 104B active. Claude Fable 5 parameter counts: unverified — no public number from this signal set.
- Context window: Kimi K3 supports 1M tokens. Fable 5 and GPT-5.6 Sol context windows: unverified — no public number from this signal set.
- Intelligence Index: Kimi K3 = 57, Fable 5 = 60, GPT-5.6 Sol = 59, per Artificial Analysis.
- Cost per Intelligence Index task: Kimi K3 = $0.94, GPT-5.6 Sol = $1.04, Fable 5 = $2.75, Opus 5 = $2.03, per Artificial Analysis.
- License: Kimi K3 is open-weight (Kimi K3 License) with weights on Hugging Face. Fable 5 and GPT-5.6 Sol are proprietary API-only.
The Reddit LLMDevs analysis makes a sharper economic point worth quoting: on one DeepSWE comparison, "Sol wins pass@1 (72.7% vs 68.5%), but Kimi is cheaper per rollout ($4.65 vs $8.37) and pulls ahead at higher pass@k." That specific number is a single third-party comparison, not a broad benchmark, but it captures the actual tradeoff builders are weighing: K3 is cheaper per attempt but often needs more attempts.

Source: @testingcatalog
Why This Matters for Builders
Kimi K3 changes the practical calculus in three ways.
First, it is the largest open-weight model ever shipped, and Moonshot committed to and delivered the open release exactly eleven days after the initial API launch. That release cadence — hosted first, weights second, in under two weeks — is now a template other labs will be measured against. It reduces the gap between "the frontier lab tried it" and "you can run it yourself" from months to days.
Second, K3's cost structure disrupts the assumption that Chinese frontier models will keep undercutting US labs on per-token sticker. Moonshot priced K3 at frontier rates because K3 is performing at near-frontier rates. As one Hacker News commenter put it on r/OpenAI: "A frontier Chinese model has no incentive to compete on price alone." The community expectation of sub-dollar frontier input tokens is now visibly wrong, at least for this generation.
Third, the open weights themselves let inference providers compete on price. Within hours of the Hugging Face drop, Mia posted a 2-bit quantization of the weights totaling 715 GB. Nebius shipped FP4 hosting on launch day. Bindu Reddy announced K3 hosted on ChatLLM and an open-source fine-tune kickoff. Expect the per-token cost of running K3 on third-party inference to fall well below Moonshot's own API pricing over the coming weeks.
If you want to run your own hands-on comparison against a comparable frontier chat model rather than rely on aggregate benchmarks, Kimi K3 is available directly on kie.ai — useful for A/B testing on your actual prompts before committing to a hosting partner.
What We Know vs. What We Don't
What we know (from official Moonshot sources and independent benchmarks):
- Kimi K3 is a 2.8T-parameter MoE with 16 of 896 experts active per token, per the Moonshot Kimi K3 blog.
- Kimi K3 supports a 1M-token context window with native vision input, per the Hugging Face model card.
- Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) with a Stable LatentMoE framework, per the official Moonshot blog.
- Kimi K3 open weights landed on Hugging Face on July 27, 2026, matching the commitment made at launch on July 16, per TestingCatalog's report.
- Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens on the official Moonshot API, with a 90% cache-hit discount, per Artificial Analysis.
- Kimi K3 launched with day-0 partners including Nebius, Baseten, Fireworks AI, DigitalOcean Serverless Inference, and Together AI, per Moonshot's launch-day announcements.
- Kimi K3 launched with Max reasoning effort only, then added Standard and High effort levels within roughly a week, per TestingCatalog.
- Kimi K3 scores 57 on the Artificial Analysis Intelligence Index, ranking fourth behind Claude Opus 5, Claude Fable 5, and GPT-5.6 Sol.
- Kimi K3 costs roughly $0.94 per Intelligence Index task versus GPT-5.6 Sol at $1.04 despite lower per-token pricing, driven by verbose output, per Artificial Analysis.
- A UK AISI / CAISI joint assessment found Kimi K3 performs significantly below frontier US models on exploit-development benchmarks, reaching step 17 of a 32-step simulated attack path versus 28.5 for leading US models, per the NIST publication.
What we don't yet know:
- Moonshot has not yet released a full Kimi K3 technical report — the launch blog promises further details on architecture, training, and evaluations in a subsequent report.
- Full-precision Kimi K3 is impractical for local hardware; 2-bit quantization brings weights to roughly 715 GB, and no smaller community quants have been publicly reported yet.
- Detailed license terms for commercial deployment (redistribution, derivative works, hosting restrictions) are governed by the "Kimi K3 License" on Hugging Face but have not been community-summarized in the current signal set.
- Third-party reproductions of Moonshot's self-reported GPU kernel optimization and MiniTriton compiler case studies have not yet appeared.
- Independent SWE-bench, Terminal-Bench, and HLE numbers for Kimi K3 from labs other than Artificial Analysis have not surfaced in the current signal set.
- The relationship between K3's training data and any distillation from Claude — a suggestion raised by one Reddit commenter noting K3's Claude-like writing style — is unverified.
How to Evaluate Kimi K3 Yourself
Given the gap between marketing benchmarks and real workloads, three practical evaluation moves are worth the time.
First, run your own cost-per-task test, not cost-per-token. Artificial Analysis's finding that K3 costs nearly the same per task as GPT-5.6 Sol despite half the sticker price is the single most important number in this release. Pick five representative tasks from your production workload, run each on K3 and your incumbent model, and log both total tokens and pass rate. The sticker-price advantage may or may not survive.
Second, test reasoning effort levels explicitly. K3 launched Max-only, and Max is the setting Artificial Analysis benchmarked. The Standard and High modes added later have not been independently benchmarked in the current signal. If your
About Elena Rossi
Elena watches developer chatter and early adoption signals to gauge which releases gain real traction.
View all posts by Elena Rossi