Skip to content

LLM API Pricing Comparison — per-token prices over time, updated monthly

Living report · updated monthly, with each Pareto edition · last updated

What a million tokens costs, tracked over time. Two data sources, both dated and verifiable: historical months reconstructed from Wayback Machine snapshots of Artificial Analysis model pages (per-model snapshot timestamps preserved in the underlying data), and the same live monthly collection that feeds the Pareto Frontier report from July 2026 on. Where no snapshot exists, the chart shows a gap — never an estimate; the archive’s coverage of AA starts February 2026. Superseded models (dashed) stay on the chart, because model churn is half the price story.

$0.3 $1 $3 $10 $30 26-0226-0326-0426-0526-0626-0726-08 Monthly snapshot (Wayback + live collection) Output $ per 1M tokens (log) Claude Fable 5 · $50 GPT-5.5 · $30 GPT-5.6 Sol · $30 Claude Opus 4.6 · $25 Claude Opus 4.7 · $25 Opus 4.8 · $25 Claude Opus 5 · $25 GPT-5.4 · $15 Kimi K3 · $15 Gemini 3.1 Pro · $12 Gemini 3 Pro · $12 Sonnet 5 · $10 Gemini 3.5 Flash · $9 GLM 5.1 · $4.4 GLM 5.2 · $4.4 Kimi K2.6 · $4 Kimi K2.7 · $4 Kimi K2.5 · $3 Gemini 3 Flash · $3 DeepSeek V3.2 · $1.6 MiniMax M2 · $1.2 MiniMax M3 · $1.2 DeepSeek V4 Flash · $0.28 MiMo V2.5 · $0.28 Open weights Proprietary Superseded (dashed, italic)

Sorted by the metric that actually matters — dollars of output per Intelligence Index point (lower = more intelligence per dollar):

ModelOutput $/MInput $/MAA Index$ per index pointWeights
DeepSeek V4 Flash$0.28$0.1450$0.006open
MiniMax M3$1.20$0.3044$0.027open
GLM 5.2$4.40$1.4051$0.086open
Kimi K2.6$4.00$0.9544$0.091open
Kimi K2.7$4.00$0.9542$0.095open
Gemini 3.5 Flash$9.00$1.5050$0.180closed
Sonnet 5$10.00$2.0053$0.189closed
Kimi K3$15.00$3.0057$0.263open
GPT-5.4$15.00$2.5051$0.294closed
Claude Opus 5$25.00$5.0061$0.410closed
Opus 4.8$25.00$5.0056$0.446closed
GPT-5.6 Sol$30.00$5.0059$0.508closed
GPT-5.5$30.00$5.0055$0.545closed
Claude Fable 5$50.00$10.0060$0.833closed
MiMo V2.5$0.28$0.14open
  • Below the flagship tier, every open model still buys intelligence cheaper than every closed model. The most efficient closed option (Gemini 3.5 Flash, $0.18 per index point) costs about twice the least efficient of the value-tier open models (Kimi K2.7, $0.095) — and 26× DeepSeek V4 Flash’s $0.007.
  • Kimi K3 broke the pattern — the first open model priced like a closed one. At $15/M output ($0.263 per index point) it sits between Sonnet 5 and GPT-5.4 on the board. What it buys: an index of 57, above everything closed except Opus 5, Fable 5 and GPT-5.6 Sol.
  • The premium is still at the top. From GLM 5.2 to Claude Opus 5 the index rises 20% (51 → 61) while the price per point rises 4.8× ($0.086 → $0.410) — and Fable 5 sits at $0.833/point. You don’t pay for intelligence — you pay for the last few points of it.
  • What to watch as editions accumulate: open-weight prices trend down because any host can serve the weights and competition compresses margins to serving cost; closed prices only move when the sole vendor decides. A year ago the best open model scored 22 on this index; today it’s 57 — though with K3, for the first time, the top open price moved up too. The value tier (GLM 5.2 and below) keeps getting smarter without getting pricier.

Break-even: when does flat-rate beat the meter?

Section titled “Break-even: when does flat-rate beat the meter?”

The per-token prices above stop mattering past a usage threshold — here it is, computed from the table (blended at the 80% input / 20% output mix typical of real workloads):

Flat subscriptionvs paying per token forBlended $/MBreak-even: tokens/monthIn agent-hours¹
Frontier, from $48.45/moGLM 5.2 itself (open, list price)$2.0024.2M~12 h/month
Frontier, from $48.45/moGPT-5.4 — same index, closed$5.009.7M~5 h/month
Frontier, from $48.45/moGPT-5.5$10.004.8M~2 h/month
Core, from $12.74/moDeepSeek V4 Flash itself (open, list price)$0.1775.8M~36 h/month
Core, from $12.74/moGPT-5.4 mini — same index, closed$1.508.5M~4 h/month

¹ We measured a coding agent at ~2.1M tokens/hour — so a single working day of agent traffic per month clears every break-even on this table against closed per-token pricing. Against the open models’ own list prices the bar is higher: about a day and a half for Frontier vs GLM 5.2, four to five days for Core vs DeepSeek’s rock-bottom rate. Below those volumes, the meter is genuinely cheaper; that’s the honest threshold.

The same thing as a picture — monthly cost against monthly usage; where a per-token diagonal crosses the flat line, the subscription starts winning:

010M20M30MTokens per month (blended 80/20 in/out)$0$50$100$150Cost per monthGPT-5.5 · $10/MGPT-5.4 · $5/MGLM 5.2 per token · $2/MFrontier flat · from $48.45/mo4.8M9.7M24.2Mflat wins →

One structural note: the flat side of this comparison is a reserved daily time block serving one request at a time per key — you’re buying capacity, not metered volume, so past break-even the marginal token costs zero for the rest of the month. The full framework for that decision is your AI bill should scale with users, not usage.

Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight models are often cheaper on aggregators, so the open rows are conservative. ”$ per index point” divides output price by a composite index — it’s a comparison heuristic, not a claim that index points are linear in value. MiMo V2.5, which we serve, appears with its current OpenRouter price but no index or history — it has no Artificial Analysis entry yet, so those cells stay empty rather than estimated.

  • 2026-07-28 — release-week collection: Kimi K3 ($3/$15), Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) join the board. Sonnet 5’s price was cut $3/$15 → $2/$10 — the first closed price drop on this tracker’s watch. K3 is the first open model priced in closed mid-tier territory.
  • 2026-07-22 — break-even table and chart recomputed on the current pool anchors (Frontier from $48.45/mo, Core from $12.74/mo with annual billing); MiMo V2.5’s OpenRouter input price rose $0.105 → $0.14/M (output unchanged at $0.28/M).
  • 2026-07-06 (backfill) — historical prices February–June 2026 reconstructed from Wayback Machine snapshots of Artificial Analysis: 19 models including superseded ones (Opus 4.6→4.7 handover, Kimi K2.5, GLM 5.1, MiniMax M2, Gemini 3 Pro/Flash). Notable finds in the record: DeepSeek V3.2’s output price rose from $0.42 to $1.60/M in May — open-weight prices mostly fall, but not always — and Kimi K2.5 wobbled $3.00 → $2.85 → $3.00.
  • 2026-07-06 (first edition) — baseline prices ingested from the 2026-07 Pareto collection. Range on the board: $0.28/M (DeepSeek V4 Flash) to $50/M (Claude Fable 5) per million output tokens — a 178× spread.

On our pools these per-token prices stop applying at all: flat monthly, no token caps during your reserved hours — Flagship from $126.65/mo, Frontier from $48.45/mo, Core from $12.74/mo. See the pools or read why your AI bill should scale with users, not usage.