Skip to content

LLM Pareto Frontier — price vs intelligence, updated monthly

Living report · updated monthly · last updated
Edition: 2026-09 (latest) · all editions: 2026-09 · 2026-08 · 2026-07

A model is Pareto-efficient if no other model is both cheaper and smarter. This report plots the current flagship LLMs — the open-weight flagships (including every one we serve) and the closed models from Anthropic, OpenAI and Google — on output price versus the Artificial Analysis Intelligence Index, and computes the frontier from the data. Everything below the line is dominated: a strictly better deal exists.

Method: all numbers come from the same source on the same day — Artificial Analysis model pages (Intelligence Index, vendor list prices). Each edition is archived so you can watch the frontier move over time.

$0$10$20$30$40$50 Output price, $ per 1M tokens 253035404550 AA Intelligence Index Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Opus 4.8 GPT-5.5 Sonnet 5 GPT-5.4 Gemini 3.5 Flash GPT-6 Astra Claude Fable 5.1 Kimi K3 Qwen3.8 Max GLM 5.2 MiniMax M3 Kimi K2.6 Kimi K2.7 DeepSeek V4 Flash GLM 5.3 GLM-5.3-Flash DeepSeek V4.1 Flash MiMo-V2.6-Pro the top open model is 7 pointsoff the top score — at 1/57 the price Open weights Proprietary Pareto frontier

Reading left to right: GLM-5.3-Flash → MiMo-V2.6-Pro → GPT-5.6 Sol → Claude Opus 5 → GPT-6 Astra → Claude Fable 5.1.

  • Every previously tracked score dropped 12–19 points this edition — that’s Artificial Analysis, not the models. AA moved its Intelligence Index to v4.3.2, a harder composite that rescored the whole field at unchanged prices (Claude Opus 5 63 → 51, GLM 5.3 60 → 45, Kimi K3 60 → 44). Scores from earlier editions are not comparable with these; compare positions, which is what the frontier is for.
  • MiMo-V2.6-Pro is the move of this edition: index 46 at $0.87/M output — the #1 open-weights model on the index, at the lowest output price of anything above 42. Xiaomi’s 1.02T-total / 42B-active MoE (MIT weights, 1M context, omnimodal, released September 21) enters the frontier at once and dominates GLM 5.3 (45 at $4.40), Kimi K3 (44 at $15), Claude Sonnet 5 (38 at $10), Gemini 3.5 Flash and DeepSeek V4.1 Flash (39 at $1.20). Full analysis: MiMo-V2.6 Pro & Flash: specs, benchmarks, pricing.
  • Two new closed models take the top of the index: GPT-6 Astra and Claude Fable 5.1, tied at 53 for $10/$50. Both join the tracked set this edition and extend the frontier upward; Claude Fable 5 (50 at the same $50) is now dominated.
  • The frontier reads GLM-5.3-Flash (42, $0.50) → MiMo-V2.6-Pro (46, $0.87) → GPT-5.6 Sol (47, $20) → Claude Opus 5 (51, $25) → GPT-6 Astra / Claude Fable 5.1 (53, $50). GLM 5.3 leaves the line — it keeps its score relative to the field (tied with Qwen3.8 Max, one point above K3) but is now beaten on both axes by MiMo-V2.6-Pro.
  • The open-vs-closed gap is 7 points (MiMo-V2.6-Pro 46 vs GPT-6 Astra and Claude Fable 5.1 at 53) on the new scale. The step from the best open model to the next closed one, GPT-5.6 Sol, is 1 point for 23× the output price — last edition’s equivalent step (GLM 5.3 → GPT-5.6 Sol) was 4.5×.
  • DeepSeek V4.1 Flash is scored at last: 39 at $1.20/M output (peak list) — five points above the V4 Flash build it replaced in our Core Pool (34 on the same scale) at a lower price, and it matches GPT-5.4 at 12.5× less. It joins the tracked set; V4 Flash stays for history.
  • Every tracked model priced under $5 per million output tokens is open-weight. GLM-5.3-Flash still anchors the cheap end ($0.50, 42) and matches Claude Opus 4.8 at 50× less. GLM 5.3, Kimi K3, Qwen3.8 Max, DeepSeek V4.1 Flash and MiMo V2.6 Flash are served in our pools; MiMo-V2.6-Pro and GLM-5.3-Flash are under evaluation (live status).
  • At the top, 2 index points (51 → 53) cost twice the output price ($25 → $50).
ModelAA IndexInput $/MOutput $/MWeightsPareto-efficient
GPT-6 Astra53$10.00$50.00closed✓
Claude Fable 5.153$10.00$50.00closed✓
Claude Opus 551$5.00$25.00closed✓
Claude Fable 550$10.00$50.00closeddominated by GPT-6 Astra
GPT-5.6 Sol47$4.00$20.00closed✓
MiMo-V2.6-Pro46$0.43$0.87open✓
GLM 5.345$1.40$4.40opendominated by MiMo-V2.6-Pro
Qwen3.8 Max45$2.00$6.00openmatched by GLM 5.3 at 1.4× less
Kimi K344$3.00$15.00opendominated by MiMo-V2.6-Pro
GLM-5.3-Flash42$0.15$0.50open✓
Opus 4.842$5.00$25.00closedmatched by GLM-5.3-Flash at 50.0× less
DeepSeek V4.1 Flash39$0.30$1.20opendominated by MiMo-V2.6-Pro
GPT-5.439$2.50$15.00closedmatched by DeepSeek V4.1 Flash at 12.5× less
Sonnet 538$2.00$10.00closeddominated by MiMo-V2.6-Pro
GPT-5.538$5.00$30.00closedmatched by Sonnet 5 at 3.0× less
DeepSeek V4 Flash34$0.44$1.32opendominated by MiMo-V2.6-Pro
GLM 5.234$1.40$4.40openmatched by DeepSeek V4 Flash at 3.3× less
Gemini 3.5 Flash33$1.50$9.00closeddominated by MiMo-V2.6-Pro
MiniMax M329$0.30$1.20opendominated by MiMo-V2.6-Pro
Kimi K2.627$0.95$4.00opendominated by MiMo-V2.6-Pro
Kimi K2.726$0.95$4.00opendominated by MiMo-V2.6-Pro

Each edition’s frontier is archived; this chart overlays them, current on top. Over time it shows the defining dynamic of this market: the frontier sliding down (cheaper) and right-side-up (smarter) — driven almost entirely by open-weight releases.

$0.00$10$20$30$40$50 Output price, $ per 1M tokens 253035404550 AA Intelligence Index 2026-09 (current) Current frontier Open Closed

First edition on Artificial Analysis Intelligence Index v4.3.2 — earlier frontiers were scored on a different scale and stay in the archived editions; previous frontiers will appear here as editions on this scale accumulate.

The Intelligence Index is one composite — task-specific rankings differ. Kimi K2.6 sits off-frontier here yet holds the best open SWE-bench Verified score (80.2); for agentic coding the ranking flips — see the Which-LLM guide for tier-fair matchups. Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight prices vary by host. MiMo V2.6 Flash, which we also serve (it replaced MiMo V2.5 in our Core Pool on 2026-09-23), has no index entry and is excluded rather than estimated — it stays on the watchlist until Artificial Analysis scores it. Index points aren’t linear in value.

  • 2026-09-23 (no frontier change) — MiMo V2.6 Flash replaces MiMo V2.5 in our Core Pool in place, at the same $0.14/$0.28 list price: 309B MoE / 15B active, image input, MIT weights. Artificial Analysis has not scored it, so it stays on the watchlist and no point on this chart moves.

  • 2026-09 (2026-09-22) — AA moved its index to v4.3.2: every previously tracked score fell 12–19 points at unchanged prices — cross-edition drops at this date are methodology, not model changes. MiMo-V2.6-Pro joins the tracked set (46, $0.43/$0.87, MIT weights) as the #1 open-weights model and enters the frontier, dominating GLM 5.3, Kimi K3, Sonnet 5 and DeepSeek V4.1 Flash — GLM 5.3 leaves the line. DeepSeek V4.1 Flash is scored (39, $0.30/$1.20 peak) and joins the tracked set. GPT-6 Astra and Claude Fable 5.1 (both 53, $10/$50) join the tracked set as the new top of the index; Claude Fable 5 becomes dominated. Frontier: GLM-5.3-Flash → MiMo-V2.6-Pro → GPT-5.6 Sol → Claude Opus 5 → GPT-6 Astra / Claude Fable 5.1. Analysis: MiMo-V2.6 Pro & Flash.

  • 2026-09-10 (no frontier change) — DeepSeek V4.1 Flash ships and replaces V4 Flash in our Core Pool: 552B MoE (8B/16B active), causal encoder–decoder, native vision, MIT weights, $0.30/$1.20 peak list — cheaper than the model it replaces. Artificial Analysis has not scored it, so it is added to the watchlist and no point on this chart moves; the DeepSeek point stays V4 Flash until an independent index exists.

  • 2026-08 refresh (2026-08-31) — GLM-5.3-Flash joins the tracked set (57, $0.15/$0.50) and lands straight on the frontier: Z.ai’s 18B-active, MIT-licensed, natively multimodal MoE matches Claude Opus 4.8 at 50× less and dominates DeepSeek V4 Flash, MiniMax M3 and Kimi K2.7 — the frontier’s whole cheap half collapses into one point. Frontier: GLM-5.3-Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5. Analysis: GLM-5.3-Flash: specs, benchmarks & pricing.

  • 2026-08 refresh (2026-08-30) — GLM 5.3 joins the tracked set (60, $1.40/$4.40) two days after its open weights shipped under the GLM-5.3 License, and enters the frontier at once: it dominates Qwen3.8 Max, Sonnet 5 and GLM 5.2, and matches Kimi K3 at 3.4× less — K3 leaves the line. GPT-5.6 Sol’s output price drops $30 → $20 and it re-enters the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5. GLM 5.2 stays tracked for history; it was upgraded in place to 5.3 in our Frontier Pool.

  • 2026-08 refresh (2026-08-14) — three moves in one collection. AA recalibrated its index to v4.1.1: every tracked score shifted up 1–5 points at unchanged prices — cross-edition score jumps around this date are methodology, not model changes. Qwen3.8 Max joins the tracked set (58, $2/$6) now that its weights shipped — it enters the frontier and pushes Sonnet 5 off it, leaving Opus 5 as the only closed model on the line. And DeepSeek V4 Flash’s peak list pricing took effect ($0.14/$0.28 → $0.44/$1.32), bringing MiniMax M3 back onto the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.

  • 2026-08 (2026-08-01) — release-week update: the V4-Flash-0731 retrain lifts DeepSeek V4 Flash 40 → 50 at unchanged prices ($0.14/$0.28) — a 10-point single-model jump. It now matches Gemini 3.5 Flash at 32× lower output price and pushes MiniMax M3 off the frontier (dominated). Frontier: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.

  • 2026-07 refresh (2026-07-28) — release-week update: Kimi K3 (57, $3/$15), Claude Opus 5 (61, $5/$25) and GPT-5.6 Sol (59, $5/$30) added to the tracked set. K3 and Opus 5 enter the frontier; Fable 5 and Opus 4.8 leave it (both dominated by Opus 5). Sonnet 5’s output price dropped $15 → $10. The open-vs-closed gap is now 4 index points — the narrowest since GLM-5 (February).

  • 2026-07 (first edition) — baseline. Frontier: DeepSeek V4 Flash, MiniMax M3, GLM 5.2, Sonnet 5, Opus 4.8, Fable 5. Notable context at launch: GLM 5.2 (released mid-June) is the top-scoring open-weights model in index history; AA’s v4.1 recalibration lowered scores across the board vs the April v4.0 figures.


GLM 5.3, Qwen3.8 Max, Kimi K3, DeepSeek V4.1 Flash and MiMo V2.6 Flash are served flat-rate in our pools — Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo — where the marginal token costs zero during your reserved hours. MiMo-V2.6-Pro and GLM-5.3-Flash are under evaluation.