LLM API Pricing Comparison — per-token prices over time, updated monthly
What a million tokens costs, tracked over time. Two data sources, both dated and verifiable: historical months reconstructed from Wayback Machine snapshots of Artificial Analysis model pages (per-model snapshot timestamps preserved in the underlying data), and the same live monthly collection that feeds the Pareto Frontier report from July 2026 on. Where no snapshot exists, the chart shows a gap — never an estimate; the archive’s coverage of AA starts February 2026. Superseded models (dashed) stay on the chart, because model churn is half the price story.
The numbers
Section titled “The numbers”Sorted by the metric that actually matters — dollars of output per Intelligence Index point (lower = more intelligence per dollar):
| Model | Output $/M | Input $/M | AA Index | $ per index point | Weights |
|---|---|---|---|---|---|
| GLM-5.3-Flash | $0.50 | $0.15 | 42 | $0.012 | open |
| MiMo-V2.6-Pro | $0.87 | $0.43 | 46 | $0.019 | open |
| DeepSeek V4.1 Flash | $1.20 | $0.30 | 39 | $0.031 | open |
| DeepSeek V4 Flash | $1.32 | $0.44 | 34 | $0.039 | open |
| MiniMax M3 | $1.20 | $0.30 | 29 | $0.041 | open |
| GLM 5.3 | $4.40 | $1.40 | 45 | $0.098 | open |
| GLM 5.2 | $4.40 | $1.40 | 34 | $0.129 | open |
| Qwen3.8 Max | $6.00 | $2.00 | 45 | $0.133 | open |
| Kimi K2.6 | $4.00 | $0.95 | 27 | $0.148 | open |
| Kimi K2.7 | $4.00 | $0.95 | 26 | $0.154 | open |
| Sonnet 5 | $10.00 | $2.00 | 38 | $0.263 | closed |
| Gemini 3.5 Flash | $9.00 | $1.50 | 33 | $0.273 | closed |
| Kimi K3 | $15.00 | $3.00 | 44 | $0.341 | open |
| GPT-5.4 | $15.00 | $2.50 | 39 | $0.385 | closed |
| GPT-5.6 Sol | $20.00 | $4.00 | 47 | $0.426 | closed |
| Claude Opus 5 | $25.00 | $5.00 | 51 | $0.490 | closed |
| Opus 4.8 | $25.00 | $5.00 | 42 | $0.595 | closed |
| GPT-5.5 | $30.00 | $5.00 | 38 | $0.789 | closed |
| GPT-6 Astra | $50.00 | $10.00 | 53 | $0.943 | closed |
| Claude Fable 5.1 | $50.00 | $10.00 | 53 | $0.943 | closed |
| Claude Fable 5 | $50.00 | $10.00 | 50 | $1.000 | closed |
| MiMo V2.5 | $0.28 | $0.14 | — | — | open |
| MiMo V2.6 Flash | $0.28 | $0.14 | — | — | open |
Analysis
Section titled “Analysis”- All $-per-point figures rose this edition because the index got harder, not because prices did. Artificial Analysis moved to Intelligence Index v4.3.2 and every score fell 12–19 points at unchanged prices. Compare ranks across editions, not the dollar figures.
- GLM-5.3-Flash still holds the floor at $0.012 per index point (42 at $0.50/M output). Right behind it is the new arrival: MiMo-V2.6-Pro, $0.019 per point — and that is the top open-weights score on the index (46) for under $1/M output, 5× cheaper per point than GLM 5.3 ($0.098) and 18× cheaper than Kimi K3. Analysis: MiMo-V2.6 Pro & Flash.
- DeepSeek V4.1 Flash is finally scored: 39 at $1.20/M peak, $0.031 per point — better on both axes than the V4 Flash build it replaced in our Core Pool (34, $0.039), and it matches GPT-5.4’s score at 12.5× lower output price.
- Below the flagship tier, every open model still buys intelligence cheaper than every closed model. The most efficient closed option (Claude Sonnet 5, $0.263 per index point) costs 1.7× the least efficient value-tier open model (Kimi K2.7, $0.154) — and 8.5× DeepSeek V4.1 Flash’s $0.031.
- Kimi K3 is still the open model priced like a closed one. At $15/M output ($0.341 per point) it sits between Gemini 3.5 Flash and GPT-5.4 on the board, for an index of 44 — above every closed model except Opus 5, Fable 5 and GPT-5.6 Sol. Qwen3.8 Max buys a point more (45) at $6/M, and MiMo-V2.6-Pro two points more (46) at $0.87/M.
- The premium is still at the top. From MiMo-V2.6-Pro to the new leaders, GPT-6 Astra and Claude Fable 5.1, the index rises 15% (46 → 53) while the price per point rises 50× ($0.019 → $0.943) — and the superseded Fable 5 sits at $1.000/point. You don’t pay for intelligence; you pay for the last few points of it.
- What to watch as editions accumulate: open-weight prices trend down because any host can serve the weights and competition compresses margins to serving cost; closed prices only move when the sole vendor decides. K3 showed the top open price can move up, and DeepSeek’s peak pricing that even the floor can rise. MiMo-V2.6-Pro is the counter-move: the top open score now costs less per output token than DeepSeek V4.1 Flash.
Break-even: when does flat-rate beat the meter?
Section titled “Break-even: when does flat-rate beat the meter?”The per-token prices above stop mattering past a usage threshold — here it is, computed from the table (blended at the 80% input / 20% output mix typical of real workloads). The full treatment — every served model, the savings factor at realistic volumes and the ceiling a single key will not go past — is its own living report: Flat Rate vs Per-Token.
| Flat subscription | vs paying per token for | Blended $/M | Break-even: tokens/month | In agent-hours¹ |
|---|---|---|---|---|
| Frontier, from $60.35/mo | GLM 5.3 itself (open, list price) | $2.00 | 30.2M | ~14 h/month |
| Frontier, from $60.35/mo | GPT-5.4 (closed) | $5.00 | 12.1M | ~6 h/month |
| Frontier, from $60.35/mo | GPT-5.5 | $10.00 | 6.0M | ~3 h/month |
| Core, from $18.70/mo | DeepSeek V4.1 Flash itself (open, peak list price) | $0.48 | 39.0M | ~19 h/month |
| Core, from $18.70/mo | GPT-5.4 — same index as DeepSeek V4.1 Flash, closed | $5.00 | 3.7M | ~1.8 h/month |
¹ We measured a coding agent at ~2.1M tokens/hour — so a single working day of agent traffic per month clears every break-even on this table against closed per-token pricing. Against the open models’ own list prices the bar is higher: about two working days for Frontier vs GLM 5.3’s list price (~14 h), and a bit more for Core vs DeepSeek V4.1 Flash’s peak rate (~19 h). Below those volumes, the meter is genuinely cheaper; that’s the honest threshold.
The same thing as a picture — monthly cost against monthly usage; where a per-token diagonal crosses the flat line, the subscription starts winning:
One structural note: the flat side of this comparison is a reserved daily time block serving one request at a time per key — you’re buying capacity, not metered volume, so past break-even the marginal token costs zero for the rest of the month. The full framework for that decision is your AI bill should scale with users, not usage.
Caveats
Section titled “Caveats”Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight models are often cheaper on aggregators, so the open rows are conservative. ”$ per index point” divides output price by a composite index — it’s a comparison heuristic, not a claim that index points are linear in value. MiMo V2.6 Flash, which we serve (it replaced MiMo V2.5 in our Core Pool on 2026-09-23), appears at Xiaomi’s list price but with no index or history — it has no Artificial Analysis entry yet, so those cells stay empty rather than estimated; MiMo V2.5 stays as a historical row at its OpenRouter price. DeepSeek V4.1 Flash is quoted at its peak rate, as AA lists it; at the off-peak rate ($0.15/$0.60) the meter is cheaper and the Core break-even roughly doubles, to ~64M tokens/month.
Changelog
Section titled “Changelog”-
2026-09-24 — Core pool block reprice for new subscriptions: blocks now $22–$26/mo (cheapest block was $17.99); annual anchor $15.29 → $18.70/mo. Existing subscriptions keep their price. Break-even rows recomputed on the new anchor.
-
2026-09-23 — MiMo V2.6 Flash replaces MiMo V2.5 in the Core Pool in place, at the same $0.14/$0.28 list price — nothing on the board moves. Both MiMo rows stay index-less until AA scores them. Model page: MiMo V2.6 Flash.
-
2026-09-22 — AA moved its Intelligence Index to v4.3.2: every previously tracked score fell 12–19 points at unchanged prices, so every $-per-point figure on the board rose — a methodology change, not a price change. MiMo-V2.6-Pro joins the board at $0.43/$0.87 (AA) with the top open-weights score (46): $0.019 per index point, second only to GLM-5.3-Flash. GPT-6 Astra and Claude Fable 5.1 (53, $10/$50 — $0.943/point) join the board as the new top closed scores. DeepSeek V4.1 Flash is scored (39 → $0.031/point) and moves from the index-less rows into the ranking; MiMo-V2.6-Flash (under review as MiMo V2.5’s successor, same $0.14/$0.28 list) takes an index-less row until AA scores it. Break-even rows re-based on DeepSeek V4.1 Flash’s peak price.
-
2026-09-10 — DeepSeek V4.1 Flash replaces V4 Flash in the Core Pool and enters the board at $0.15/$0.60 off-peak, $0.30/$1.20 peak — down from V4 Flash’s $0.22/$0.66 and $0.44/$1.32. Cache hits are listed at $0.003/$0.006 per 1M. No AA index yet, so no $-per-point cell and no change to the Pareto frontier; V4 Flash stays on the board as the historical row. What changed.
-
2026-08-31 — GLM-5.3-Flash joins the board and takes the top row outright: $0.009 per index point (57 at $0.15/$0.50 list, MIT weights) — 2.8× below DeepSeek V4 Flash’s $0.025, the previous best, with near-frontier intelligence rather than value-tier. Z.ai lists a 50% launch promo on top until September 9.
-
2026-08-30 — GPT-5.6 Sol cut its list price $5/$30 → $4/$20: the first closed price cut since Sonnet 5’s in July, and enough to re-enter the Pareto frontier at 61.
-
2026-08-27 — Core pool block reprice: blocks now $17.99–$21.99/mo (cheapest block was $16.49); annual anchor $14.02 → $15.29/mo. Break-even rows recomputed on the new anchor.
-
2026-08-14 — DeepSeek V4 Flash $0.14/$0.28 → $0.44/$1.32: the peak/off-peak pricing DeepSeek announced for August 16 is now what Artificial Analysis lists (peak rate shown; off-peak stays near the old level). Qwen3.8 Max joins the board (58, $2/$6 — $0.103/point) now that its weights are open. All AA index scores shifted up 1–5 points with AA’s v4.1.1 recalibration; the $-per-point column is recomputed on the new scores.
-
2026-08-14 — Core pool block prices updated ($16.49–$21.99/mo per block; annual anchor $12.74 → $14.02/mo); break-even rows recomputed. DeepSeek announced peak/off-peak API pricing effective Aug 16 (V4 Flash output $0.28 → $0.66–$1.32/M) — the per-token board picks it up in the next monthly collection.
-
2026-07-28 — release-week collection: Kimi K3 ($3/$15), Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) join the board. Sonnet 5’s price was cut $3/$15 → $2/$10 — the first closed price drop on this tracker’s watch. K3 is the first open model priced in closed mid-tier territory.
-
2026-07-22 — break-even table and chart recomputed on the pool anchors current at the time (Frontier from $60.35/mo, Core at $12.74/mo with annual billing — since raised to $14.02/mo, see the 2026-08-14 entry); MiMo V2.5’s OpenRouter input price rose $0.105 → $0.14/M (output unchanged at $0.28/M).
-
2026-07-06 (backfill) — historical prices February–June 2026 reconstructed from Wayback Machine snapshots of Artificial Analysis: 19 models including superseded ones (Opus 4.6→4.7 handover, Kimi K2.5, GLM 5.1, MiniMax M2, Gemini 3 Pro/Flash). Notable finds in the record: DeepSeek V3.2’s output price rose from $0.42 to $1.60/M in May — open-weight prices mostly fall, but not always — and Kimi K2.5 wobbled $3.00 → $2.85 → $3.00.
-
2026-07-06 (first edition) — baseline prices ingested from the 2026-07 Pareto collection. Range on the board: $0.28/M (DeepSeek V4 Flash) to $50/M (Claude Fable 5) per million output tokens — a 178× spread.
On our pools these per-token prices stop applying at all: flat monthly, no token caps during your reserved hours — Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo. See the pools or read why your AI bill should scale with users, not usage.