Skip to content

LLM Pareto Frontier — price vs intelligence, updated monthly

Living report · updated monthly · last updated
Edition: 2026-08 (latest) · all editions: 2026-08 · 2026-07

A model is Pareto-efficient if no other model is both cheaper and smarter. This report plots the current flagship LLMs — the open-weight models we serve and the closed models from Anthropic, OpenAI and Google — on output price versus the Artificial Analysis Intelligence Index, and computes the frontier from the data. Everything below the line is dominated: a strictly better deal exists.

Method: all numbers come from the same source on the same day — Artificial Analysis model pages (Intelligence Index, vendor list prices). Each edition is archived so you can watch the frontier move over time.

$0$10$20$30$40$50 Output price, $ per 1M tokens 4045505560 AA Intelligence Index Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Opus 4.8 GPT-5.5 Sonnet 5 GPT-5.4 Gemini 3.5 Flash Kimi K3 GLM 5.2 MiniMax M3 Kimi K2.6 Kimi K2.7 DeepSeek V4 Flash the top open model is now 4 pointsoff the top score — for 40% less Open weights (we serve them) Proprietary Pareto frontier

Reading left to right: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.

  • Kimi K3 lands 4 index points off the top score (57 vs Claude Opus 5’s 61) — per Artificial Analysis, the narrowest open-vs-closed gap since the GLM-5 release in February — and it does it at $15/M output vs Opus 5’s $25.
  • Claude Fable 5 fell off the frontier. At 60 it is dominated by its own stablemate Opus 5: one point smarter at half the output price. GPT-5.6 Sol (59, $30) arrives already dominated too.
  • Below $10 per million output tokens, the frontier is entirely open-weight (DeepSeek V4 Flash, GLM 5.2) — and every open model on the frontier, K3 included, is in our pools.
  • The V4-Flash-0731 retrain is this edition’s story: +10 index points (40 → 50) at unchanged prices moved DeepSeek V4 Flash into Gemini 3.5 Flash territory at 32× less per output token — and knocked MiniMax M3 off the frontier entirely (now dominated by a cheaper, smarter model).
  • The near-vertical cliff at the top is gone: the last step used to be GLM 5.2 (51) → Fable 5 (60) at 11.4× the price; it is now Kimi K3 (57) → Opus 5 (61), 4 points for 1.7×.
  • Sonnet 5’s quiet price cut ($15 → $10 output) keeps it on the frontier as the only non-open model besides Opus 5.
ModelAA IndexInput $/MOutput $/MWeightsPareto-efficient
Claude Opus 561$5.00$25.00closed
Claude Fable 560$10.00$50.00closeddominated by Claude Opus 5
GPT-5.6 Sol59$5.00$30.00closeddominated by Claude Opus 5
Kimi K357$3.00$15.00open
Opus 4.856$5.00$25.00closeddominated by Claude Opus 5
GPT-5.555$5.00$30.00closeddominated by Claude Opus 5
Sonnet 553$2.00$10.00closed
GLM 5.251$1.40$4.40open
GPT-5.451$2.50$15.00closedmatched by GLM 5.2 at 3.4× less
DeepSeek V4 Flash50$0.14$0.28open
Gemini 3.5 Flash50$1.50$9.00closedmatched by DeepSeek V4 Flash at 32.1× less
MiniMax M344$0.30$1.20opendominated by DeepSeek V4 Flash
Kimi K2.644$0.95$4.00openmatched by MiniMax M3 at 3.3× less
Kimi K2.742$0.95$4.00opendominated by DeepSeek V4 Flash

Each edition’s frontier is archived; this chart overlays them, current on top. Over time it shows the defining dynamic of this market: the frontier sliding down (cheaper) and right-side-up (smarter) — driven almost entirely by open-weight releases.

$0.00$10$20$30$40$50 Output price, $ per 1M tokens 4045505560 AA Intelligence Index 2026-07 2026-08 (current) Current frontier Previous editions (older = fainter) Open Closed

The Intelligence Index is one composite — task-specific rankings differ. Kimi K2.6 sits off-frontier here yet holds the best open SWE-bench Verified score (80.2); for agentic coding the ranking flips — see the Which-LLM guide for tier-fair matchups. Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight prices vary by host. MiMo V2.5, which we also serve, has no index entry yet and is excluded rather than estimated. Index points aren’t linear in value.

  • 2026-08 (2026-08-01) — release-week update: the V4-Flash-0731 retrain lifts DeepSeek V4 Flash 40 → 50 at unchanged prices ($0.14/$0.28) — a 10-point single-model jump. It now matches Gemini 3.5 Flash at 32× lower output price and pushes MiniMax M3 off the frontier (dominated). Frontier: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.

  • 2026-07 refresh (2026-07-28) — release-week update: Kimi K3 (57, $3/$15), Claude Opus 5 (61, $5/$25) and GPT-5.6 Sol (59, $5/$30) added to the tracked set. K3 and Opus 5 enter the frontier; Fable 5 and Opus 4.8 leave it (both dominated by Opus 5). Sonnet 5’s output price dropped $15 → $10. The open-vs-closed gap is now 4 index points — the narrowest since GLM-5 (February).

  • 2026-07 (first edition) — baseline. Frontier: DeepSeek V4 Flash, MiniMax M3, GLM 5.2, Sonnet 5, Opus 4.8, Fable 5. Notable context at launch: GLM 5.2 (released mid-June) is the top-scoring open-weights model in index history; AA’s v4.1 recalibration lowered scores across the board vs the April v4.0 figures.


All four open-weight models on the frontier — Kimi K3 included — are served flat-rate in our pools — Flagship from $126.65/mo, Frontier from $48.45/mo, Core from $12.74/mo — where the marginal token costs zero during your reserved hours.