LLM Pareto Frontier — price vs intelligence, updated monthly
A model is Pareto-efficient if no other model is both cheaper and smarter. This report plots the current flagship LLMs — the open-weight flagships (including every one we serve) and the closed models from Anthropic, OpenAI and Google — on output price versus the Artificial Analysis Intelligence Index, and computes the frontier from the data. Everything below the line is dominated: a strictly better deal exists.
Method: all numbers come from the same source on the same day — Artificial Analysis model pages (Intelligence Index, vendor list prices). Each edition is archived so you can watch the frontier move over time.
Current frontier (edition 2026-09)
Section titled “Current frontier (edition 2026-09)”Reading left to right: GLM-5.3-Flash → MiMo-V2.6-Pro → GPT-5.6 Sol → Claude Opus 5 → GPT-6 Astra → Claude Fable 5.1.
Analysis
Section titled “Analysis”- Every previously tracked score dropped 12–19 points this edition — that’s Artificial Analysis, not the models. AA moved its Intelligence Index to v4.3.2, a harder composite that rescored the whole field at unchanged prices (Claude Opus 5 63 → 51, GLM 5.3 60 → 45, Kimi K3 60 → 44). Scores from earlier editions are not comparable with these; compare positions, which is what the frontier is for.
- MiMo-V2.6-Pro is the move of this edition: index 46 at $0.87/M output — the #1 open-weights model on the index, at the lowest output price of anything above 42. Xiaomi’s 1.02T-total / 42B-active MoE (MIT weights, 1M context, omnimodal, released September 21) enters the frontier at once and dominates GLM 5.3 (45 at $4.40), Kimi K3 (44 at $15), Claude Sonnet 5 (38 at $10), Gemini 3.5 Flash and DeepSeek V4.1 Flash (39 at $1.20). Full analysis: MiMo-V2.6 Pro & Flash: specs, benchmarks, pricing.
- Two new closed models take the top of the index: GPT-6 Astra and Claude Fable 5.1, tied at 53 for $10/$50. Both join the tracked set this edition and extend the frontier upward; Claude Fable 5 (50 at the same $50) is now dominated.
- The frontier reads GLM-5.3-Flash (42, $0.50) → MiMo-V2.6-Pro (46, $0.87) → GPT-5.6 Sol (47, $20) → Claude Opus 5 (51, $25) → GPT-6 Astra / Claude Fable 5.1 (53, $50). GLM 5.3 leaves the line — it keeps its score relative to the field (tied with Qwen3.8 Max, one point above K3) but is now beaten on both axes by MiMo-V2.6-Pro.
- The open-vs-closed gap is 7 points (MiMo-V2.6-Pro 46 vs GPT-6 Astra and Claude Fable 5.1 at 53) on the new scale. The step from the best open model to the next closed one, GPT-5.6 Sol, is 1 point for 23× the output price — last edition’s equivalent step (GLM 5.3 → GPT-5.6 Sol) was 4.5×.
- DeepSeek V4.1 Flash is scored at last: 39 at $1.20/M output (peak list) — five points above the V4 Flash build it replaced in our Core Pool (34 on the same scale) at a lower price, and it matches GPT-5.4 at 12.5× less. It joins the tracked set; V4 Flash stays for history.
- Every tracked model priced under $5 per million output tokens is open-weight. GLM-5.3-Flash still anchors the cheap end ($0.50, 42) and matches Claude Opus 4.8 at 50× less. GLM 5.3, Kimi K3, Qwen3.8 Max, DeepSeek V4.1 Flash and MiMo V2.6 Flash are served in our pools; MiMo-V2.6-Pro and GLM-5.3-Flash are under evaluation (live status).
- At the top, 2 index points (51 → 53) cost twice the output price ($25 → $50).
The data
Section titled “The data”| Model | AA Index | Input $/M | Output $/M | Weights | Pareto-efficient |
|---|---|---|---|---|---|
| GPT-6 Astra | 53 | $10.00 | $50.00 | closed | ✓ |
| Claude Fable 5.1 | 53 | $10.00 | $50.00 | closed | ✓ |
| Claude Opus 5 | 51 | $5.00 | $25.00 | closed | ✓ |
| Claude Fable 5 | 50 | $10.00 | $50.00 | closed | dominated by GPT-6 Astra |
| GPT-5.6 Sol | 47 | $4.00 | $20.00 | closed | ✓ |
| MiMo-V2.6-Pro | 46 | $0.43 | $0.87 | open | ✓ |
| GLM 5.3 | 45 | $1.40 | $4.40 | open | dominated by MiMo-V2.6-Pro |
| Qwen3.8 Max | 45 | $2.00 | $6.00 | open | matched by GLM 5.3 at 1.4× less |
| Kimi K3 | 44 | $3.00 | $15.00 | open | dominated by MiMo-V2.6-Pro |
| GLM-5.3-Flash | 42 | $0.15 | $0.50 | open | ✓ |
| Opus 4.8 | 42 | $5.00 | $25.00 | closed | matched by GLM-5.3-Flash at 50.0× less |
| DeepSeek V4.1 Flash | 39 | $0.30 | $1.20 | open | dominated by MiMo-V2.6-Pro |
| GPT-5.4 | 39 | $2.50 | $15.00 | closed | matched by DeepSeek V4.1 Flash at 12.5× less |
| Sonnet 5 | 38 | $2.00 | $10.00 | closed | dominated by MiMo-V2.6-Pro |
| GPT-5.5 | 38 | $5.00 | $30.00 | closed | matched by Sonnet 5 at 3.0× less |
| DeepSeek V4 Flash | 34 | $0.44 | $1.32 | open | dominated by MiMo-V2.6-Pro |
| GLM 5.2 | 34 | $1.40 | $4.40 | open | matched by DeepSeek V4 Flash at 3.3× less |
| Gemini 3.5 Flash | 33 | $1.50 | $9.00 | closed | dominated by MiMo-V2.6-Pro |
| MiniMax M3 | 29 | $0.30 | $1.20 | open | dominated by MiMo-V2.6-Pro |
| Kimi K2.6 | 27 | $0.95 | $4.00 | open | dominated by MiMo-V2.6-Pro |
| Kimi K2.7 | 26 | $0.95 | $4.00 | open | dominated by MiMo-V2.6-Pro |
How the frontier moves
Section titled “How the frontier moves”Each edition’s frontier is archived; this chart overlays them, current on top. Over time it shows the defining dynamic of this market: the frontier sliding down (cheaper) and right-side-up (smarter) — driven almost entirely by open-weight releases.
First edition on Artificial Analysis Intelligence Index v4.3.2 — earlier frontiers were scored on a different scale and stay in the archived editions; previous frontiers will appear here as editions on this scale accumulate.
Caveats
Section titled “Caveats”The Intelligence Index is one composite — task-specific rankings differ. Kimi K2.6 sits off-frontier here yet holds the best open SWE-bench Verified score (80.2); for agentic coding the ranking flips — see the Which-LLM guide for tier-fair matchups. Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight prices vary by host. MiMo V2.6 Flash, which we also serve (it replaced MiMo V2.5 in our Core Pool on 2026-09-23), has no index entry and is excluded rather than estimated — it stays on the watchlist until Artificial Analysis scores it. Index points aren’t linear in value.
Changelog
Section titled “Changelog”-
2026-09-23 (no frontier change) — MiMo V2.6 Flash replaces MiMo V2.5 in our Core Pool in place, at the same $0.14/$0.28 list price: 309B MoE / 15B active, image input, MIT weights. Artificial Analysis has not scored it, so it stays on the watchlist and no point on this chart moves.
-
2026-09 (2026-09-22) — AA moved its index to v4.3.2: every previously tracked score fell 12–19 points at unchanged prices — cross-edition drops at this date are methodology, not model changes. MiMo-V2.6-Pro joins the tracked set (46, $0.43/$0.87, MIT weights) as the #1 open-weights model and enters the frontier, dominating GLM 5.3, Kimi K3, Sonnet 5 and DeepSeek V4.1 Flash — GLM 5.3 leaves the line. DeepSeek V4.1 Flash is scored (39, $0.30/$1.20 peak) and joins the tracked set. GPT-6 Astra and Claude Fable 5.1 (both 53, $10/$50) join the tracked set as the new top of the index; Claude Fable 5 becomes dominated. Frontier: GLM-5.3-Flash → MiMo-V2.6-Pro → GPT-5.6 Sol → Claude Opus 5 → GPT-6 Astra / Claude Fable 5.1. Analysis: MiMo-V2.6 Pro & Flash.
-
2026-09-10 (no frontier change) — DeepSeek V4.1 Flash ships and replaces V4 Flash in our Core Pool: 552B MoE (8B/16B active), causal encoder–decoder, native vision, MIT weights, $0.30/$1.20 peak list — cheaper than the model it replaces. Artificial Analysis has not scored it, so it is added to the watchlist and no point on this chart moves; the DeepSeek point stays V4 Flash until an independent index exists.
-
2026-08 refresh (2026-08-31) — GLM-5.3-Flash joins the tracked set (57, $0.15/$0.50) and lands straight on the frontier: Z.ai’s 18B-active, MIT-licensed, natively multimodal MoE matches Claude Opus 4.8 at 50× less and dominates DeepSeek V4 Flash, MiniMax M3 and Kimi K2.7 — the frontier’s whole cheap half collapses into one point. Frontier: GLM-5.3-Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5. Analysis: GLM-5.3-Flash: specs, benchmarks & pricing.
-
2026-08 refresh (2026-08-30) — GLM 5.3 joins the tracked set (60, $1.40/$4.40) two days after its open weights shipped under the GLM-5.3 License, and enters the frontier at once: it dominates Qwen3.8 Max, Sonnet 5 and GLM 5.2, and matches Kimi K3 at 3.4× less — K3 leaves the line. GPT-5.6 Sol’s output price drops $30 → $20 and it re-enters the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5. GLM 5.2 stays tracked for history; it was upgraded in place to 5.3 in our Frontier Pool.
-
2026-08 refresh (2026-08-14) — three moves in one collection. AA recalibrated its index to v4.1.1: every tracked score shifted up 1–5 points at unchanged prices — cross-edition score jumps around this date are methodology, not model changes. Qwen3.8 Max joins the tracked set (58, $2/$6) now that its weights shipped — it enters the frontier and pushes Sonnet 5 off it, leaving Opus 5 as the only closed model on the line. And DeepSeek V4 Flash’s peak list pricing took effect ($0.14/$0.28 → $0.44/$1.32), bringing MiniMax M3 back onto the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.
-
2026-08 (2026-08-01) — release-week update: the V4-Flash-0731 retrain lifts DeepSeek V4 Flash 40 → 50 at unchanged prices ($0.14/$0.28) — a 10-point single-model jump. It now matches Gemini 3.5 Flash at 32× lower output price and pushes MiniMax M3 off the frontier (dominated). Frontier: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.
-
2026-07 refresh (2026-07-28) — release-week update: Kimi K3 (57, $3/$15), Claude Opus 5 (61, $5/$25) and GPT-5.6 Sol (59, $5/$30) added to the tracked set. K3 and Opus 5 enter the frontier; Fable 5 and Opus 4.8 leave it (both dominated by Opus 5). Sonnet 5’s output price dropped $15 → $10. The open-vs-closed gap is now 4 index points — the narrowest since GLM-5 (February).
-
2026-07 (first edition) — baseline. Frontier: DeepSeek V4 Flash, MiniMax M3, GLM 5.2, Sonnet 5, Opus 4.8, Fable 5. Notable context at launch: GLM 5.2 (released mid-June) is the top-scoring open-weights model in index history; AA’s v4.1 recalibration lowered scores across the board vs the April v4.0 figures.
GLM 5.3, Qwen3.8 Max, Kimi K3, DeepSeek V4.1 Flash and MiMo V2.6 Flash are served flat-rate in our pools — Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo — where the marginal token costs zero during your reserved hours. MiMo-V2.6-Pro and GLM-5.3-Flash are under evaluation.