State of Open Weights — the open models we serve, and why
Open a model router today and you’ll find 400+ models — the same family in five quantizations, three deprecation states, a dozen providers with different latencies and silent fallbacks. Choice isn’t the product; it’s a tax. This report is the opposite: the small set of open-weight models we actually operate, each one there for a job, with the facts that matter — real license terms, architecture, context, current prices and scores — kept current so you never have to reconcile stale numbers across blog posts.
Sources: model facts from the official Hugging Face cards (linked per row); prices and Intelligence Index from Artificial Analysis, refreshed monthly by the same collector that updates the Pareto Frontier report.
The lineup
Section titled “The lineup”| Model | Pool | The job | Params | Context | License |
|---|---|---|---|---|---|
| Kimi K3 | Flagship | Open flagship — 9 points off the top closed score on the AA index (44, v4.3.2) | 2.8T MoE / 104B active | 1M | Kimi K3 License³ |
| Qwen3.8 Max | Flagship | Second flagship profile — hybrid on-demand reasoning, vision input | 2.4T MoE / 95B active | 262K native / ~1M extended | Qwen3.8-Max License⁵ |
| Kimi K2.7 | — (retired August 2026) | Agentic coding, tool use at a fraction of K3’s price | 1T MoE / 32B active | 256K | Modified MIT¹ |
| Kimi K2.6 | — (retired July 2026) | Best open SWE-bench Verified score; pinned agent workflows | 1T MoE / 32B active | 256K | Modified MIT¹ |
| GLM 5.3 | Frontier | Long-horizon coding; 45 on the AA index (v4.3.2) at $4.40/M output | 753B MoE | 1M | GLM-5.3 License⁴ |
| GLM-5.3-Flash | — (under evaluation) | Near-frontier at the price floor — the cheapest point on the Pareto frontier; native multimodal input | 320B MoE / 18B active | 1M | MIT |
| GLM 5.2 | — (upgraded to GLM 5.3, August 2026) | Long-horizon coding; superseded in place by GLM 5.3 (same base, extended post-training) | 753B MoE | 1M | MIT |
| MiniMax M3 | Frontier | 1M context + native multimodal input | ~428B MoE / 23B active | 1M | Community License² |
| DeepSeek V4.1 Flash | Core | High-volume workloads: extraction, chat, summarization, RAG — with native image input | 552B MoE / 8B active (prefill), 16B (decode) | 1M | MIT |
| DeepSeek V4 Flash | — (upgraded to V4.1 Flash, September 2026) | High-volume workloads; superseded in place by DeepSeek V4.1 Flash (new generation) | 284B MoE / 13B active | 1M | MIT |
| MiMo V2.5 | — (upgraded to V2.6 Flash, September 2026) | Budget multimodal; superseded in place by MiMo V2.6 Flash (new generation, same size class and list price) | 310B MoE / 15B active | 1M | MIT |
| MiMo V2.6 Flash | Core | Budget multimodal workhorse — vendor agent benchmarks within 4 points of MiMo-V2.6-Pro at the same list price as MiMo V2.5; image input | 309B MoE / 15B active | 1M | MIT |
| MiMo-V2.6-Pro | — (under evaluation) | Top open-weights score on the AA index (46, v4.3.2) at $0.87/M output; omnimodal input | 1.02T MoE / 42B active | 1M | MIT |
¹ Modified MIT: attribution required only above 100M MAU or $20M/mo revenue. ² MiniMax Community License: “Built with MiniMax M3” attribution; >$20M/yr revenue requires written authorization. ³ Kimi K3 License: Moonshot’s own license for the K3 weights — read the license text before self-hosting. ⁴ GLM-5.3 License: Z.ai’s custom license (GLM 5.2 was MIT, and so is GLM-5.3-Flash) — commercial use with attribution; a Model-as-a-Service operator with more than US$10B aggregate revenue over 12 months must pass Z.ai’s security review first — license text. ⁵ Qwen3.8-Max License: custom Alibaba license — attribution at consumer scale; paid license for MaaS/assistant products above $50M revenue. “Open” is not one thing — MIT and Apache 2.0 mean unrestricted commercial use; community licenses sit between open and proprietary. For API consumers none of this matters day to day; it matters if you later self-host or embed weights in a product.
Current numbers
Section titled “Current numbers”| Model | Input $/M | Output $/M | AA Index | Headline benchmark |
|---|---|---|---|---|
| Kimi K3 | $3 | $15 | 44 | Terminal-Bench 2.1 88.3 · GPQA Diamond 93.5 |
| Qwen3.8 Max | $2 | $6 | 45 | SWE-bench Pro 67.7 · Terminal-Bench 2.1 86.6 |
| Kimi K2.7 | $0.95 | $4 | 26 | MCP Mark Verified 81.1 (vendor suite) |
| Kimi K2.6 | $0.95 | $4 | 27 | SWE-bench Verified 80.2 · Pro 58.6 |
| GLM 5.3 | $1.4 | $4.4 | 45 | Terminal-Bench 2.1 88.2 · SWE-Marathon 42.5 (vendor) |
| GLM-5.3-Flash | $0.15 | $0.5 | 42 | Terminal-Bench 2.1 84.3 · DeepSWE 63.4 (vendor) |
| GLM 5.2 | $1.4 | $4.4 | 34 | SWE-bench Pro 62.1 · Terminal-Bench 2.1 81.0 |
| MiniMax M3 | $0.3 | $1.2 | 29 | SWE-bench Verified 80.5 · Pro 59.0 |
| DeepSeek V4.1 Flash | $0.3 | $1.2 | 39 | Terminal-Bench 2.1 90.6 · DeepSWE v1.1 74.2 (vendor) |
| DeepSeek V4 Flash | $0.44 | $1.32 | 34 | Terminal-Bench 2.1: 82.7 (0731, vendor) |
| MiMo V2.5 | $0.14 | $0.28 | — | SWE-bench Pro 56.1 · Terminal-Bench 2 65.8 |
| MiMo-V2.6-Pro | $0.43 | $0.87 | 46 | Terminal-Bench 2.1 89.9 · DeepSWE v1.1 71.9 (vendor) |
| MiMo V2.6 Flash | $0.14 | $0.28 | — | Terminal-Bench 2.1 87.6 · DeepSWE v1.1 67.9 (vendor) |
Prices are the vendors’ list prices per million tokens as tracked by Artificial Analysis (DeepSeek at its peak rate; MiMo V2.5: OpenRouter listing; MiMo V2.6 Flash: Xiaomi’s list price — neither has an AA index yet). Index scores are AA’s v4.3.2 edition, which rescored every model 12–19 points lower than v4.1.1; don’t compare them with older figures. Benchmark scores are as published on each model’s official card; scaffolding differs between labs. Where these models sit against the closed flagships — and which are Pareto-efficient — is the Pareto Frontier report; tier-fair matchups are in the Which-LLM guide.
How to choose between the pools
Section titled “How to choose between the pools”Flagship when you want the strongest open models we serve — Kimi K3 and Qwen3.8 Max, 9 and 8 points off the top closed score on the current AA index. Frontier when the work is agentic coding, multi-step agents, or anything where a failed run costs more than the tokens did. Core when the work is volume — extraction, classification, summarization, chat, pipelines that run all day. All three are flat-rate: no token caps during your reserved hours, so the per-token prices above stop mattering once you’re inside your window.
Changelog
Section titled “Changelog”-
2026-09-23 — MiMo V2.6 Flash joins the roster in the Core Pool, replacing MiMo V2.5 in place: same size class (309B MoE / 15B active), same $0.14/$0.28 list price, MIT weights (XiaomiMiMo/MiMo-V2.6-Flash-RL), image input, and a generation of RL post-training — vendor agent benchmarks within 4 points of the 1T MiMo-V2.6-Pro. Artificial Analysis has not scored it yet, so its index cell stays empty rather than estimated and MiMo V2.5 keeps its own row for history. Old id
mimo-v2.5is served until 2026-10-23. Model page: MiMo V2.6 Flash; analysis: MiMo-V2.6 Pro & Flash. -
2026-09-22 — MiMo-V2.6 added to the tracked set: Xiaomi released both models with MIT weights on September 21. MiMo-V2.6-Pro (1.02T MoE / 42B active, under evaluation) debuts as the top open-weights model on the AA index (46) at $0.43/$0.87. MiMo-V2.6-Flash (309B / 15B active) is under review as MiMo V2.5’s successor in the Core Pool, at the same list price. DeepSeek V4.1 Flash is scored (39). Board-wide: AA moved its index to v4.3.2, and every score fell 12–19 points at unchanged prices (K3 44, GLM 5.3 and Qwen3.8 Max 45) — methodology, not the models. Analysis: MiMo-V2.6 Pro & Flash.
-
2026-09-10 — DeepSeek V4.1 Flash joins the roster in the Core Pool, replacing V4 Flash in place: a new generation rather than a retrain — 552B MoE (8B active in prefill, 16B in decode) on a new causal encoder–decoder architecture, vision trained in from pre-training, 1M context, MIT weights (deepseek-ai/DeepSeek-V4.1-Flash) — and a lower list price than the model it replaces ($0.15/$0.60 off-peak, $0.30/$1.20 peak, down from $0.22/$0.66 and $0.44/$1.32). Artificial Analysis has not scored it yet, so its index cell stays empty rather than estimated and V4 Flash keeps its own row for history. Old model id
deepseek-v4-flashis accepted until 2026-10-10. Analysis: DeepSeek V4.1 Flash: what changed. -
2026-08-31 — GLM-5.3-Flash added to the tracked set (under evaluation — not in a pool): Z.ai’s 320B MoE / 18B-active sibling to GLM 5.3, natively multimodal, 1M context, plain-MIT weights (zai-org/GLM-5.3-Flash) — 57 on the AA index at $0.15/$0.50 list, straight onto the Pareto frontier. Qwen3.8 Max also gains its missing lineup row (served in Flagship since 08-14). Analysis: GLM-5.3-Flash: specs, benchmarks & pricing.
-
2026-08-30 — GLM 5.3 joins the roster in the Frontier Pool, replacing GLM 5.2 in place: Z.ai published the weights (zai-org/GLM-5.3, 753B MoE, fp8) on August 28 under the custom GLM-5.3 License — not MIT like 5.2 — and the model scores 60 on the AA index (v4.1.1), tied with Kimi K3 at $4.40/M output. GLM 5.2 stays in the table for history.
-
2026-08-14 — Qwen3.8 Max joins the roster in the Flagship Pool: its open weights (Qwen3.8-2.4T-A95B, 2.4T MoE / 95B active) shipped August 12 under the custom Qwen3.8-Max License, lifting the licensing gate that held it in review. Kimi K2.7 retired from the Frontier Pool in the same window. Board-wide: AA’s Intelligence Index recalibration to v4.1.1 lifted every score 1–5 points at unchanged prices (K3 60, Qwen 58, GLM 5.2 53), and DeepSeek V4 Flash’s list price moved to its new peak rate ($0.44/$1.32).
-
2026-08-01 — DeepSeek V4 Flash re-scored 40 → 50 on the AA index at unchanged prices, following the V4-Flash-0731 retrain.
-
2026-07-28 — Kimi K3 joins the roster in the new Flagship Pool: 2.8T MoE (104B active), 1M context, index 57 — 4 points off the top closed model, the narrowest open-vs-closed gap since February per Artificial Analysis. Kimi K2.6’s index nudged 43 → 44 in the same collection.
-
2026-07-27 — Kimi K2.6 retired from the Frontier Pool; Kimi K2.7 remains its successor in the same pool.
-
2026-07-22 — MiMo V2.5 input price on its OpenRouter listing rose $0.105 → $0.14/M; output unchanged at $0.28/M. Still no Artificial Analysis entry.
-
2026-07-06 (first edition) — baseline roster: Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3 (Frontier); DeepSeek V4 Flash, MiMo V2.5 (Core). This report absorbs the per-model tables previously maintained in blog posts, which now link here instead of carrying their own copies of the numbers.
The live pool menu — with current block prices and availability — is always at cheapestinference.com/pools: Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo, one OpenAI- and Anthropic-compatible API.