Kimi K3 vs Claude: 2.8T Open Model vs Opus 4.8

Daniel Okonkwo

Daniel Okonkwo

Senior ML Engineer

Published: July 15, 2026
Kimi K3 vs Claude Opus 4.8 side-by-side comparison chart

TLDRKimi K3 launched at 2.8T params + 1M context as open weights; Claude Opus 4.8 is Anthropic's shipping flagship. Full spec, price, and use-case comparison.

Kimi K3 vs Claude: 2.8T-Param Open Challenger Meets Opus 4.8

Kimi K3 is Moonshot AI's next-generation open-weight model, launched July 16, 2026 with a confirmed 2.8T-parameter architecture, a 1M-token context window, and a long-horizon agent focus, while Claude (Anthropic's Opus 4.8 flagship) is a closed-weight, production-hardened frontier model with confirmed pricing and top independent coding benchmarks — the right pick now depends on whether you need open weights and a bigger context window (K3) or verified multi-agent reliability and mature tooling (Claude). Independent head-to-head benchmark data is still thin: launch-day community tests place K3 near Opus 4.8 on coding and design, but third-party SWE-Bench and Terminal-Bench numbers have yet to land.

Updated 2026-07-16: Kimi K3 is now live on Kimi at a confirmed 2.8T parameters with a 1M context, per Moonshot's official platform docs and Financial Times reporting (see the Update below).

Key Takeaways

  • Kimi K3 launched on July 16, 2026 with a confirmed 2.8T total parameters, a 1M-token context, Kimi Delta Attention (hybrid linear) + Attention Residuals architecture, native visual understanding, and always-on thinking mode.
  • Claude Opus 4.8 is Anthropic's shipping flagship, with Opus 4.7 confirmed at 87.6% on SWE-Bench Verified and 91/100 on Kilo Code's FlowGraph workflow test.
  • Pricing gap has narrowed sharply: K3 launched at $3.00 input / $0.30 cache hit / $15.00 output per million tokens — roughly 1:1 with Anthropic's Sonnet series, a break from the ~10x discount K2.6 held over Opus 4.7.
  • Open weights vs closed: K3 is positioned as open-weight following Moonshot's Modified MIT precedent, with weights expected on Hugging Face in the coming days; Claude is closed and API-only.
  • Coding verdict is tightening: launch-day testers report K3 at Opus 4.8 level on coding with strong frontend design taste; Fable 5 and GPT-5.6 Sol still lead on Terminal-Bench 2.1.
  • Long-horizon agents are the stated K3 focus, with K3 Swarm Max shipping alongside K3 Max for large-scale parallel processing, extending K2.6's 300 sub-agent Agent Swarm.

Kimi K3 vs Claude at a Glance

DimensionKimi K3 (Moonshot AI)Claude Opus 4.8 (Anthropic)
StatusLaunched July 16, 2026; live on Kimi web, iOS app, Kimi Code, APIShipping; current Claude flagship
ArchitectureMoE, 2.8T total params (official); Kimi Delta Attention (hybrid linear) + Attention ResidualsNot publicly disclosed
Context window1,048,576 tokens (1M) — officialNot detailed in this signal bundle
MultimodalNative visual understanding (text + image in → text out)Text and vision (per prior Claude generations)
License / weightsOpen-weight (weights expected in coming days; Modified MIT precedent from K2)Closed, API-only
Confirmed coding benchmarkLaunch-day: "Opus 4.8 level in coding" per community teardowns; independent SWE-Bench pendingOpus 4.7: 87.6% SWE-Bench Verified
Reference pricing$3.00 input / $0.30 cache hit / $15.00 output per MTok (official)Not detailed in this signal bundle; roughly Anthropic Sonnet-tier parity
Long-horizon runsK3 Swarm Max variant for large-scale parallel processing; K2.6 precedent of 12+ hour agent tasksNot detailed in this signal bundle
AvailabilityKimi web app, iOS app, Kimi Code, Moonshot API (kimi-k3)Anthropic API, Claude apps

Architecture and Capabilities

Kimi K3 is Moonshot's most capable model to date at 2.8 trillion total parameters, built on Kimi Delta Attention — a hybrid linear attention mechanism — plus Attention Residuals, with native visual understanding and a 1M-token context window per Moonshot's official platform docs. Max Weinbach flagged the practical serving question after launch: "Kimi K3 being 2.8T parameter and 1M context is cool but show me the sparsity, show me the price. How quick can the inference providers scale this to 200 tok/s" (source). The full sparsity ratio and active-parameter count have not yet been published.

Claude Opus 4.8 is Anthropic's current flagship. Anthropic has not published a full architecture disclosure, and this signal bundle does not contain official Opus 4.8 spec detail. What is known from prior generations: Claude is closed-weight, accessible only through Anthropic's API and partner platforms, and Opus 4.7 was the direct predecessor benchmarked against Kimi K2.6.

The philosophical split matters. Kimi's K2.6 release demonstrated 13-hour autonomous coding runs with 1,000+ tool calls and 300 sub-agents via its "Agent Swarm" system (Medium writeup). K3 extends that trajectory with K3 Swarm Max, a variant explicitly built for large-scale parallel processing that shipped alongside K3 Max on launch day (source). Claude is optimized for a different curve — reliability inside shorter, tool-heavy loops that production teams can trust.

For the deep-dive spec context, see our earlier writeup on what Kimi K3 is.

Benchmarks

Independent third-party benchmark data for Kimi K3 is still landing. The most useful pre-launch signal came from Chatbot Arena's stealth model "Kivine," widely identified by testers as the K3 preview. TestingCatalog captured a side-by-side of Kivine (K3) against Claude Fable 5 on a universe-simulation prompt: "Fable 5 finished faster, and most UX components were more robust and easy to use. Kimi K3 was much more complex and visually appealing" (source).

Post-launch, community testers have converged on an "Opus 4.8 level in coding" read for K3, with ChrissGPT summarizing the specs at rollout as "2-3T Parms, Open Source, Opus 4.8 level in coding" (source). Teortaxes' pre-launch estimate of ~55 on Artificial Analysis — a band he placed near Opus 4.8 and GPT-5.5 — remains the working expectation until independent scoring lands (source).

For Claude, the last verified generation (Opus 4.7) has hard numbers:

TestClaude Opus 4.7Kimi K2.6 (prior K-family)
SWE-Bench Verified87.6%80.2%
Kilo Code FlowGraph workflow91/10068/100

These are per the Medium writeup on K2.6, which noted the "23-point gap concentrated in lease handling, cross-run scheduling, and live SSE streaming — exactly the kind of multi-agent contention bugs that don't show up in benchmark suites." Whether K3's new Kimi Delta Attention architecture closes that gap is the open question.

Pricing

This axis has changed more than any other with the K3 launch. Kimi K2.6, the direct predecessor, launched at $0.60 input / $2.50 output per million tokens — described as "roughly 8.3× cheaper on input and 10× cheaper on output than Claude Opus 4.7" in third-party analysis.

Kimi K3's official API pricing is now published: $3.00 per million input tokens (cache miss), $0.30 per million on cache hits, and $15.00 per million output tokens, with flat rates across the full 1M-token context window. That's roughly 1:1 with Anthropic's Sonnet series and close to GPT-5.6 Terra pricing — a dramatic step up from the K2.6 discount, and roughly half the per-token cost of GPT-5.6 Sol at $30/M output. Alongside the API launch, Moonshot began a Chinese-market top-up promotion with bonus credits ranging from 10% (¥99–¥499) up to 30% (¥5,000+), running July 15 through August 11 (source).

Kimi K3 launch promo page from Moonshot's Kimi API platform

Source: @LuminaXspace

Claude's Opus 4.8 rate card is not in this signal bundle. Historically Anthropic prices Opus at premium tiers with Sonnet and Haiku at progressive discounts, and there is no indication that pattern changes with 4.8. The upshot: K3 no longer wins on price at the flagship tier the way K2.6 did — Moonshot's automatic context caching (which drops input to $0.30/M for stable prefixes) is where the meaningful discount now lives.

Availability and Access

Claude is available today through Anthropic's official API, Claude apps, and partner integrations. Opus 4.8 access is straightforward for anyone with an Anthropic account.

Kimi K3 went live on July 16, 2026. As of launch day:

  • Both K3 Max (chat and agent tasks) and K3 Swarm Max (large-scale parallel processing) are available on Kimi's web app and iOS app (source, source).
  • Kimi Code ships K3 as an available model for developers (source).
  • The Moonshot API exposes the model as kimi-k3 at https://api.moonshot.ai/v1, with thinking mode always enabled via reasoning_effort=max.
  • Open weights on Hugging Face are expected in the coming days but had not landed as of launch day.

Update — 2026-07-16

Moonshot began rolling out Kimi K3 on July 16, 2026, hours after officially teasing the model on X (source). Both K3 Max and K3 Swarm Max variants are live inside the Kimi app and Kimi Code, with users confirming access within the first hour of launch (source, source). Financial Times reporting cited by community accounts pegs the model at 2–3T parameters, a 1M-token context, and performance targeted to exceed Claude Opus 4.8 (source).

The parameter count has since sharpened to a specific 2.8T total, per community teardowns of the released model — putting K3 roughly on par with the reported 3T-parameter footprint of Claude Fable 5 and GPT-5.6 Sol, and materially above the 1.5T figure attributed to Opus 4.8 (source, source). One early oddity worth noting for prompt-tuning teams: K3's chain-of-thought output is reportedly rendered in English regardless of the input language (source).

Two caveats temper the hype. Sparsity, per-token pricing, and inference throughput are still unpublished — the axes Max Weinbach flagged as the ones that actually determine whether a 2.8T model is deployable at scale (source). And independent Artificial Analysis scoring has not landed; Teortaxes' pre-release estimate of "~55 on Artificial Analysis" — a band he places near Opus 4.8 and GPT-5.5 — remains the working expectation rather than a verified number (source). SWE-Bench Verified and Terminal-Bench 2.1 results from third-party evaluators are the next shoe to drop.

Building similar chat and reasoning workloads? On kie.ai you can try Kimi K3, Claude Opus 5, and Claude Sonnet 5.

For related release-timing context, see our coverage of the DeepSeek V4 release and Anthropic's next-generation Claude Opus 5.

Which One Should You Use?

Choose Kimi K3 if:

  • You need open weights for on-premises deployment, fine-tuning, or air-gapped environments (once weights land on Hugging Face in the coming days).
  • Your workflow benefits from Moonshot's automatic context caching — the $0.30/M cache-hit input rate rewards stable long prompts and tool definitions.
  • You're building long-horizon agent workflows that can take advantage of K3 Swarm Max for large-scale parallel processing.
  • A 1M-token context is a hard requirement for your use case.
  • You want to be on the newest frontier open model and can tolerate independent benchmarks arriving over the coming weeks.

Choose Claude if:

  • Your workload is coding-heavy and you value the verified 87.6% SWE-Bench Verified ceiling on Opus 4.7 (with 4.8 expected higher).
  • You depend on mature multi-agent contention handling — the area where independent tests showed the K2.6-vs-Opus-4.7 gap was largest.
  • Closed-weight is acceptable and you don't need to self-host.
  • You're integrating with tools that already have first-class Anthropic support.
  • Price parity at the flagship tier makes the reliability premium worth it for your workload.

Frequently Asked Questions

Is Kimi K3 better than Claude? Independent third-party benchmarks are still landing, but early community testing after the July 16, 2026 launch places Kimi K3 near Claude Opus 4.8 and GPT-5.5 in the ~55 Artificial Analysis band, with Financial Times reporting positioning it above Opus 4.8 on mainstream benchmarks while still trailing Fable 5 overall. Claude Opus 4.7 remains the last generation with fully verified numbers like SWE-Bench Verified at 87.6%.

Is Kimi K3 cheaper than Claude? Not by much at the flagship tier. Moonshot's official API pricing for Kimi K3 is $3.00 input (cache miss) / $0.30 cache hit / $15.00 output per million tokens — roughly 1:1 with Anthropic's Sonnet series and only about half the per-token cost of GPT-5.6 Sol. That's a break from the ~10x discount Kimi K2.6 held over Claude Opus 4.7 at $0.60 / $2.50.

Which is better for coding, Kimi K3 or Claude? Early launch-day tests place Kimi K3 at Claude Opus 4.8 level on coding, with community demos showing it trading blows with GPT-5.6 Sol on tasks like voxel rendering and frontend design. Claude Opus 4.7 still holds the strongest verified coding record (87.6% SWE-Bench Verified, 91/100 Kilo Code FlowGraph), and independent SWE-Bench and Terminal-Bench 2.1 results for K3 have not yet been published.

Does Kimi K3 have a longer context window than Claude? Yes for most Claude tiers. Kimi K3 ships with a confirmed 1,048,576-token (1M) context window per Moonshot's official platform docs. Claude Opus 4.8's context length is not detailed in this signal bundle, though Claude Sonnet 4.6 was reported at 1M in beta earlier in 2

Daniel Okonkwo

About Daniel Okonkwo

Daniel writes about inference systems, model architecture, and what new releases actually change for builders.

View all posts by Daniel Okonkwo