Kimi K3: specs, benchmarks, and our day-one plan to serve it
Kimi K3 is Moonshot AI’s new flagship, announced on July 16 after a leaked promotion page on Moonshot’s own platform tipped the release a day early. The headline is simple: on the independent Artificial Analysis Intelligence Index it scores 57 — above Claude Opus 4.8 (56), making it the first open-weight model to outscore a Claude Opus-class model on that index.
And to answer the question this blog exists to answer: yes, we will serve it. We have already secured the capacity to run a model of this size. K3 goes live on CheapestInference the moment two things are true: the weights are actually published, and the license terms permit commercial serving. Nothing else is in the way.
Update — July 2026: Kimi K3 is now live. The weights shipped on schedule (July 27), and K3 is served in the new Flagship Pool with unlimited usage at a flat monthly price — very limited seats. How to use the Kimi K3 API → · Subscribe →
July 28: with Claude Opus 5 (61) and GPT-5.6 Sol (59) landing the same month, K3’s 57 puts the open-vs-closed gap at 4 index points — per Artificial Analysis, the narrowest since the GLM-5 release in February. Only two labs score higher. Our Pareto Frontier report now charts K3 — straight onto the frontier.
What Moonshot announced
Section titled “What Moonshot announced”| Architecture | Mixture-of-Experts, ~2.8T total parameters |
| Context window | 1M tokens |
| Input | Text, image, and video |
| Variants at launch | K3 Max (chat and agent tasks) · K3 Swarm Max (large-scale parallel processing) |
| Available today | Moonshot’s API, Kimi Code, and the Kimi app |
| Open weights | Published July 27, 2026 — Hugging Face, Kimi K3 License |
| List price (API) | $3 input / $15 output per 1M tokens |
Coverage from the launch day: TechCrunch on the Opus gap closing, Fortune on Chinese AI entering Fable-level territory, and Simon Willison’s notes for a practitioner’s first look.
When is the Kimi K3 release date?
Section titled “When is the Kimi K3 release date?”K3 was announced on July 16, 2026, usable from day one through Moonshot’s own API, Kimi Code, and the Kimi app. The open weights were published on July 27, 2026, on schedule, on Hugging Face under Moonshot’s own Kimi K3 License — and K3 went live on CheapestInference’s Flagship Pool the same week.
The numbers so far
Section titled “The numbers so far”- Artificial Analysis Intelligence Index: 57. For scale: Claude Opus 4.8 scores 56, and the best open-weight model until now — GLM 5.2 — scores 51: the open ceiling jumped six points in one release. Only three closed models score higher — Claude Opus 5 (61), Claude Fable 5 (60) and GPT-5.6 Sol (59). Where every model sits on price-vs-intelligence is our Pareto Frontier report, whose refreshed 2026-07 edition now charts K3 — straight onto the Pareto frontier.
- #1 on Frontend Code Arena with 1,679 points — ahead of Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and GLM 5.2 (1,587).
- The Kimi track record. Moonshot’s K2.6 still holds the best open SWE-bench Verified score (80.2), and the K2 line has been the default open choice for tool-heavy agent work — see the tier-fair matchups in our Which-LLM guide.
The honest caveats, same as in our living reports: the index is one composite and task-specific rankings differ, and these are launch-week numbers, mostly on Moonshot’s own serving stack. The weights question resolved on schedule: K3 is downloadable from Hugging Face under Moonshot’s own Kimi K3 License (not the K2 line’s Modified MIT — read the text before self-hosting), and K3 now has its row in State of Open Weights.
What it changed
Section titled “What it changed”When K3 was announced, every model above GLM 5.2’s intelligence score was closed, and the open-vs-closed gap had sat at 5+ index points all year. With the weights shipped, the open ceiling is 57 — and after Claude Opus 5 (61) and GPT-5.6 Sol (59) landed in the same month, the gap to the very top stands at 4 index points, the narrowest since the GLM-5 release in February (Artificial Analysis). K3’s $3/$15 list price still undercuts every closed model in its class, and it walked straight onto the Pareto frontier.
Kimi K3 API pricing
Section titled “Kimi K3 API pricing”Moonshot’s list price for the K3 API is $3 input / $15 output per 1M tokens — undercutting every closed model in its class, but 3–4× the K2 line’s price: the first open flagship priced like a closed mid-tier model (Price Tracker). On CheapestInference that call is made: K3 debuted in its own Flagship Pool, flat-rate from $149/mo, unlimited usage during your reserved hours and very limited seats — the pools page is always the live source for lineup, prices and availability.
Unlimited Kimi K3 API access
Section titled “Unlimited Kimi K3 API access”- Live now. The capacity we secured before launch is serving K3 today in the Flagship Pool — we didn’t start the clock the day the weights dropped.
- The usual: one OpenAI- and Anthropic-compatible API, flat-rate time-block subscriptions, no token caps during your reserved hours — so it drops into Claude Code, Cline, or any compatible client; it’s in
GET /v1/modelsnow.
If you want to be running K3 this week, create an account — seats in the Flagship Pool are very limited, no waitlist.
CheapestInference serves Kimi K3 (Flagship Pool, from $149/mo), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool, from $48.45/mo billed annually) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool, from $12.74/mo) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.