Moonshot AI's Kimi K3 is the largest open-weight model ever shipped. We break down how it stacks up against the closed US flagships it's chasing — OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 — on benchmarks, price, and openness.
Moonshot AI just dropped Kimi K3, and it's the biggest open-weight model the world has seen: a 2.8-trillion-parameter Mixture-of-Experts system with the new Kimi Delta Attention mechanism and a 1M-token context window. The weights go public by July 27, 2026 — which is the real headline. An open model this capable, this cheap, changes the math for anyone building on top of frontier AI.
But "biggest open model" isn't the same as "best model." So the question everyone's actually asking: how does K3 stack up against the two closed flagships it's chasing — OpenAI's GPT‑5.6 Sol and Anthropic's Claude Fable 5?
The short version
Kimi K3 leapfrogs last generation's frontier (it beats Claude Opus 4.8 and GPT‑5.5 on Moonshot's own numbers) but still trails the current US flagships — Fable 5 and GPT‑5.6 Sol — on the hardest reasoning and agentic work. Where it wins is the combination almost nobody else offers: near-frontier quality, open weights, and aggressive pricing.
How they compare
| Kimi K3 | GPT‑5.6 Sol | Claude Fable 5 | |
|---|---|---|---|
| Vendor | Moonshot AI | OpenAI | Anthropic |
| Released | Jul 16, 2026 | GA Jul 9, 2026 | Jun 9, 2026 |
| Openness | Open weights (by Jul 27) | Closed / API | Closed / API |
| Context | 1M tokens | 1.1M tokens | 1M tokens |
| Price (in / out per 1M) | $3 / $15 | $5 / $30 | $10 / $50 |
| Coding (Terminal‑Bench 2.1) | 88.3% | 88.8% (Sol Ultra 91.9%) | — |
| SWE‑Bench Pro | — | — | 80.3% |
| Reasoning (GPQA Diamond) | 93.5% | — | — |
| Agentic browsing (BrowseComp) | 91.2% | — | — |
| LMArena Frontend Code | #1, 1679 Elo | — | — |
Benchmarks are drawn from each lab's launch materials and public trackers; cross-vendor tests don't always use identical harnesses, so treat exact deltas as directional, not gospel.
Where Kimi K3 wins
- Price-to-performance. At $3/$15, K3 undercuts Sol (~40% cheaper) and Fable 5 (~70% cheaper output) while landing in the same conversation on coding and reasoning. For high-volume workloads, that gap compounds fast.
- Open weights. This is the differentiator. You can self-host, fine-tune, and run K3 in air-gapped or data-sovereign environments — something neither Sol nor Fable 5 allows at any price.
- Front-end code generation. K3 debuted at #1 on LMArena's Frontend Code Arena (1679 Elo) — genuinely best-in-class for UI/component generation right now.
- Raw knowledge & browsing. 93.5% on GPQA Diamond and 91.2% on BrowseComp are top-tier at release.
Where Sol and Fable 5 still lead
- Fable 5 remains the reasoning and software-engineering benchmark to beat — 80.3% on SWE‑Bench Pro and a 1932 GDPval‑AA score, ahead of everything else Anthropic has published. It's the pick when correctness on hard, multi-step engineering matters more than cost. (Worth knowing: Fable 5 ships with safety classifiers that hand certain cyber/bio/chem requests off to Opus 4.8 mid-response.)
- GPT‑5.6 Sol edges K3 on agentic coding — Sol Ultra hits 91.9% on Terminal‑Bench 2.1 vs K3's 88.3% — and the Sol/Terra/Luna tiering lets you dial cost vs. capability per task. Sol also carries the deepest tool/ecosystem integration (ChatGPT, Codex, the API).
The bottom line — which one, when?
- Reach for Kimi K3 when you want frontier-adjacent quality at the lowest cost, need open weights (self-hosting, fine-tuning, data residency), or you're generating a lot of front-end code.
- Reach for GPT‑5.6 Sol for top-end agentic coding and mature tooling, with cost-tiering flexibility.
- Reach for Claude Fable 5 when hard reasoning and software-engineering accuracy are non-negotiable and budget is secondary.
The real story of K3 isn't that it beats the US flagships — it doesn't, quite. It's that the gap between open and closed frontier models is now measured in single-digit benchmark points, and the open side costs a fraction as much. That's a very different landscape than we had even six months ago.
Sources: MarkTechPost · VentureBeat · Simon Willison · llm-stats (GPT‑5.6 Sol) · The Agent Report · Anthropic — Fable 5 & Mythos 5 · Vellum