State of Open Weights — the open models we serve, and why
Open a model router today and you’ll find 400+ models — the same family in five quantizations, three deprecation states, a dozen providers with different latencies and silent fallbacks. Choice isn’t the product; it’s a tax. This report is the opposite: the small set of open-weight models we actually operate, each one there for a job, with the facts that matter — real license terms, architecture, context, current prices and scores — kept current so you never have to reconcile stale numbers across blog posts.
Sources: model facts from the official Hugging Face cards (linked per row); prices and Intelligence Index from Artificial Analysis, refreshed monthly by the same collector that updates the Pareto Frontier report.
The lineup
Section titled “The lineup”| Model | Pool | The job | Params | Context | License |
|---|---|---|---|---|---|
| Kimi K3 | Flagship | Open flagship — 4 points off the top closed model on the AA index | 2.8T MoE / 104B active | 1M | Kimi K3 License³ |
| Kimi K2.7 | Frontier | Agentic coding, tool use at a fraction of K3’s price | 1T MoE / 32B active | 256K | Modified MIT¹ |
| Kimi K2.6 | — (retired July 2026) | Best open SWE-bench Verified score; pinned agent workflows | 1T MoE / 32B active | 256K | Modified MIT¹ |
| GLM 5.2 | Frontier | Long-horizon coding; best open intelligence-per-dollar on the frontier | 753B MoE | 1M | MIT |
| MiniMax M3 | Frontier | 1M context + native multimodal input | ~428B MoE / 23B active | 1M | Community License² |
| DeepSeek V4 Flash | Core | High-volume workloads: extraction, chat, summarization, RAG | 284B MoE / 13B active | 1M | MIT |
| MiMo V2.5 | Core | Budget multimodal; strongest coding card in its price class | 310B MoE / 15B active | 1M | MIT |
¹ Modified MIT: attribution required only above 100M MAU or $20M/mo revenue. ² MiniMax Community License: “Built with MiniMax M3” attribution; >$20M/yr revenue requires written authorization. ³ Kimi K3 License: Moonshot’s own license for the K3 weights — read the license text before self-hosting. “Open” is not one thing — MIT and Apache 2.0 mean unrestricted commercial use; community licenses sit between open and proprietary. For API consumers none of this matters day to day; it matters if you later self-host or embed weights in a product.
Current numbers
Section titled “Current numbers”| Model | Input $/M | Output $/M | AA Index | Headline benchmark |
|---|---|---|---|---|
| Kimi K3 | $3 | $15 | 57 | Terminal-Bench 2.1 88.3 · GPQA Diamond 93.5 |
| Kimi K2.7 | $0.95 | $4 | 42 | MCP Mark Verified 81.1 (vendor suite) |
| Kimi K2.6 | $0.95 | $4 | 44 | SWE-bench Verified 80.2 · Pro 58.6 |
| GLM 5.2 | $1.4 | $4.4 | 51 | SWE-bench Pro 62.1 · Terminal-Bench 2.1 81.0 |
| MiniMax M3 | $0.3 | $1.2 | 44 | SWE-bench Verified 80.5 · Pro 59.0 |
| DeepSeek V4 Flash | $0.14 | $0.28 | 50 | LiveCodeBench 91.6 |
| MiMo V2.5 | $0.14 | $0.28 | — | SWE-bench Pro 56.1 · Terminal-Bench 2 65.8 |
Prices are the vendors’ list prices per million tokens as tracked by Artificial Analysis (MiMo V2.5: OpenRouter listing — no AA entry yet). Benchmark scores are as published on each model’s official card; scaffolding differs between labs. Where these models sit against the closed flagships — and which are Pareto-efficient — is the Pareto Frontier report; tier-fair matchups are in the Which-LLM guide.
How to choose between the pools
Section titled “How to choose between the pools”Flagship when you want the strongest open model there is — Kimi K3, 4 points off the top closed score. Frontier when the work is agentic coding, multi-step agents, or anything where a failed run costs more than the tokens did. Core when the work is volume — extraction, classification, summarization, chat, pipelines that run all day. All three are flat-rate: no token caps during your reserved hours, so the per-token prices above stop mattering once you’re inside your window.
Changelog
Section titled “Changelog”-
2026-08-01 (auto-collected draft — review) — DeepSeek V4 Flash: $0.14/$0.28 idx 40 → $0.14/$0.28 idx 50.
-
2026-07-28 — Kimi K3 joins the roster in the new Flagship Pool: 2.8T MoE (104B active), 1M context, index 57 — 4 points off the top closed model, the narrowest open-vs-closed gap since February per Artificial Analysis. Kimi K2.6’s index nudged 43 → 44 in the same collection.
-
2026-07-27 — Kimi K2.6 retired from the Frontier Pool; Kimi K2.7 remains its successor in the same pool.
-
2026-07-22 — MiMo V2.5 input price on its OpenRouter listing rose $0.105 → $0.14/M; output unchanged at $0.28/M. Still no Artificial Analysis entry.
-
2026-07-06 (first edition) — baseline roster: Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3 (Frontier); DeepSeek V4 Flash, MiMo V2.5 (Core). This report absorbs the per-model tables previously maintained in blog posts, which now link here instead of carrying their own copies of the numbers.
The live pool menu — with current block prices and availability — is always at cheapestinference.com/pools: Flagship from $126.65/mo, Frontier from $48.45/mo, Core from $12.74/mo, one OpenAI- and Anthropic-compatible API.