Which LLM should you use? A living buyer's guide
Most model comparisons make a category error: they put a $0.28-per-million model and a $50-per-million model in the same list and act surprised when one “wins”. This guide works the way the decision actually works — three steps: pick your weight class, pick open or closed within it, pick the model. Numbers refresh monthly from the same collection as the Pareto Frontier report; every score links to its source.
Step 1 — Pick your weight class
Section titled “Step 1 — Pick your weight class”| Your workload | Class |
|---|---|
| Agentic coding, multi-step agents, complex reasoning — anywhere a failed run costs more than the tokens did | Heavyweight |
| Extraction, classification, summarization, chat, RAG, pipelines that run all day | Budget |
The classes are real: closed heavyweights cost up to 180× more per output token than budget models, and nothing in the budget class should be benchmarked against Claude Opus.
Step 2a — Heavyweight: open or closed?
Section titled “Step 2a — Heavyweight: open or closed?”| Model | AA Index | SWE-bench Verified | SWE-bench Pro | Output $/M | Weights |
|---|---|---|---|---|---|
| GPT-6 Astra | 53 | — | — | $50.00 | closed |
| Claude Fable 5.1 | 53 | — | — | $50.00 | closed |
| Claude Opus 5 | 51 | — | — | $25.00 | closed |
| Claude Fable 5 | 50 | ~95.0 | — | $50.00 | closed |
| GPT-5.6 Sol | 47 | — | — | $20.00 | closed |
| MiMo-V2.6-Pro | 46 | — | — | $0.87 | open |
| GLM 5.3 | 45 | — | — | $4.40 | open |
| Qwen3.8 Max | 45 | — | 67.7 | $6.00 | open |
| Kimi K3 | 44 | — | — | $15.00 | open |
| Claude Opus 4.8 | 42 | 88.6 | 69.2 | $25.00 | closed |
| Claude Sonnet 5 | 38 | — | — | $10.00 | closed |
| GPT-5.5 | 38 | — | 58.6 | $30.00 | closed |
| GLM 5.2 | 34 | — | 62.1 | $4.40 | open |
| Gemini 3.1 Pro | 30 | 80.6 | — | $12.00 | closed |
| MiniMax M3 | 29 | 80.5 | 59.0 | $1.20 | open |
| Kimi K2.6 | 27 | 80.2 | 58.6 | $4.00 | open |
| Kimi K2.7 | 26 | — | — | $4.00 | open |
Choose closed when you need the ceiling. The very top is still closed — GPT-6 Astra and Claude Fable 5.1 (53) and Claude Opus 5 (51) on the index, 88.6–95 on SWE-bench Verified for Opus 4.8 and Fable 5, and on long-horizon agentic reliability (METR time horizons) every frontier entry is closed. The best open model, MiMo-V2.6-Pro, sits 7 index points below the top score on AA’s current v4.3.2 scale. If the last few points pay for themselves, that fight is closed — at $20–50 per million output tokens.
Choose open for everything below the ceiling — which now starts one rung from the top. MiMo-V2.6-Pro (46, at $0.87/M), GLM 5.3 and Qwen3.8 Max (45, at $4.40/M and $6/M) and Kimi K3 (44, at $15/M) all outscore Opus 4.8, GPT-5.5 and Sonnet 5 on the index — and MiMo-V2.6-Pro does it for less per output token than any closed model on the budget table below. GLM 5.2 already outscored GPT-5.5 on SWE-bench Pro (62.1 vs 58.6) at one-seventh the price; MiniMax M3 and Kimi K2.6 sit at statistical parity with Gemini 3.1 Pro on Verified at one-tenth and one-third of its price. And open weights carry strategic advantages no benchmark shows:
- No deprecation risk. Closed models get retired on the vendor’s schedule — OpenAI retired GPT-4o in February 2026; Anthropic’s deprecation page lists six Claude retirements in nine months. Downloaded weights cannot be retired; the model you validated is the model you run.
- Control. No closed vendor trains on paid API data by default, but with open weights policy isn’t the only protection: the same model can move to any host or your own hardware. Licenses vary — see the fine print per model in State of Open Weights.
- A competitive market. Every closed model has exactly one seller; open weights are served by competing hosts, which is why their prices keep falling (Price Tracker) and why flat-rate offers exist at all.
The Kimi paradox, or why the index misleads here: Kimi K2.6 scores 27 on the composite index (v4.3.2), near the foot of the table above, yet holds the best open SWE-bench Verified score (80.2) — it’s an agentic-coding specialist, and Artificial Analysis called it “the new leading open weights model” at launch. Usage agrees: Chinese-origin open models carry over 45% of OpenRouter’s token volume, and the heaviest token burners are coding agents. When a composite index and the market disagree, look at what the market does with the model.
Step 2b — Budget: open or closed?
Section titled “Step 2b — Budget: open or closed?”| Model | AA Index | Input $/M | Output $/M | Weights |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 39 | $0.30 | $1.20 | open |
| Gemini 3.5 Flash | 33 | $1.50 | $9.00 | closed |
| GPT-5.4 mini | 24 | $0.75 | $4.50 | closed |
| Gemini 3 Flash* | 18 | $0.50 | $3.00 | closed |
| GPT-5 mini | 17 | $0.25 | $2.00 | closed |
| Claude 4.5 Haiku* | 15 | $1.00 | $5.00 | closed |
| MiMo V2.5 | — | $0.14 | $0.28 | open |
| MiMo V2.6 Flash | — | $0.14 | $0.28 | open |
* Non-reasoning variants as evaluated by Artificial Analysis; the others are reasoning variants. DeepSeek V4.1 Flash is quoted at its peak list rate. MiMo V2.5’s price is its OpenRouter listing and MiMo V2.6 Flash’s is Xiaomi’s list price — neither has an AA index yet, so those cells stay empty rather than estimated.
This division isn’t close on today’s data: DeepSeek V4.1 Flash beats GPT-5.4 mini by 15 index points (39 vs 24) at 3.75× less per output token, and every other closed model under $1 input scores 21+ points below it. The one strong closed entry, Gemini 3.5 Flash (33), costs $9/M output — twice the price of GLM 5.3 (45), a heavyweight. When a vendor’s budget model costs more than the other side’s heavyweight, the weight classes have inverted. MiMo V2.6 Flash — MiMo V2.5’s in-place successor in our Core Pool since 2026-09-23, at the same $0.28/M — adds image input and reports vendor agent benchmarks within 4 points of the 1T MiMo-V2.6-Pro (67.9 DeepSWE v1.1, 87.6 Terminal-Bench 2.1), against a 56.1 SWE-bench Pro card score for the model it replaced. Z.ai’s GLM-5.3-Flash (42 on the index at $0.50/M output, MIT weights) still tops the Price Tracker’s dollars-per-index-point board.
Step 3 — Which model for which job
Section titled “Step 3 — Which model for which job”| The job | Reach for | Why |
|---|---|---|
| Frontier-class work on open weights | Kimi K3 or Qwen3.8 Max | Index 44–45 — the flagship open profiles, 6–7 points off the top closed score |
| Agentic coding, tool-heavy agents | Kimi K2.7 / K2.6 | Best open SWE-bench Verified (80.2); K2.7 is the token-efficiency refresh |
| Long-horizon coding on a budget | GLM 5.3 | Index 45 at $4.40/M — ties Qwen3.8 Max at 1.4× less |
| Whole-codebase / long-document / multimodal | MiniMax M3 | 1M context, image+video input, ~80 SWE-bench Verified at $1.20/M |
| High-volume pipelines, chat, extraction | DeepSeek V4.1 Flash | The budget workhorse — index 39 (GPT-5.4’s score) at $1.20/M output peak list; 1M context, native image input, MIT weights |
| Budget multimodal | MiMo V2.6 Flash | Image input at the lowest price on the board; vendor agent scores within 4 points of the 1T MiMo-V2.6-Pro |
| The absolute capability ceiling | GPT-6 Astra / Claude Fable 5.1 / Claude Opus 5 | Index 53/53/51 — when the last points pay for themselves |
Kimi K3, Qwen3.8 Max, GLM 5.3, MiniMax M3, DeepSeek V4.1 Flash and MiMo V2.6 Flash are served in our pools flat-rate — Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo, no token caps during your reserved hours.
Head-to-head: the pairs people actually compare
Section titled “Head-to-head: the pairs people actually compare”Search traffic clusters on a few specific match-ups. Here is each one, with the same sourced numbers as the tables above — they refresh monthly from the same collection, so nothing below is a frozen figure. The short version first: most “X vs Y” questions are really which job, not which is better. Where a vendor publishes no score for a cell, it shows — rather than an estimate.
Kimi K3 vs Kimi K2.7
Section titled “Kimi K3 vs Kimi K2.7”| Model | AA Index | SWE-bench Verified | SWE-bench Pro | Input $/M | Output $/M | Context | Weights |
|---|---|---|---|---|---|---|---|
| Kimi K3 | 44 | — | — | $3.00 | $15.00 | 1M | open |
| Kimi K2.7 | 26 | — | — | $0.95 | $4.00 | 256K | open |
| GLM 5.3 | 45 | — | — | $1.40 | $4.40 | 1M | open |
Same lab, different questions. K3 is Moonshot’s open flagship — 18 index points over K2.7, a 1M context (vs 256K) and frontier-class scores, at roughly 4× the per-token list price. K2.7 remains the agentic-coding workhorse: tuned for tool-heavy runs at $4/M. If the task is hard enough that you were considering a closed frontier model, K3 is now the open answer; if the job is steady agent loops where cost compounds, K2.7 (or GLM 5.3, the intelligence-per-dollar pick) still wins the invoice.
Kimi vs DeepSeek
Section titled “Kimi vs DeepSeek”| Model | AA Index | SWE-bench Verified | SWE-bench Pro | Input $/M | Output $/M | Context | Weights |
|---|---|---|---|---|---|---|---|
| Kimi K2.7 | 26 | — | — | $0.95 | $4.00 | 256K | open |
| Kimi K2.6 | 27 | 80.2 | 58.6 | $0.95 | $4.00 | 256K | open |
| DeepSeek V4.1 Flash | 39 | — | — | $0.30 | $1.20 | 1M | open |
Different weight classes doing different jobs, not rivals for the same task. Kimi is the agentic-coding specialist — reach for it when a failed multi-step, tool-heavy run costs more than the tokens it burns; K2.6 holds the benchmark score and K2.7 is the token-efficiency refresh. DeepSeek V4.1 Flash is the budget workhorse: high-volume extraction, chat, RAG and pipelines that run all day, at a fraction of Kimi’s per-token cost. Pick Kimi for the hard agent loops, DeepSeek when throughput and price dominate.
DeepSeek V4.1 Flash is served flat-rate in our pools — no per-token metering during your reserved hours. Kimi K2.7 was retired in August 2026; the Moonshot model we serve today is Kimi K3 (Flagship Pool).
GLM vs Kimi
Section titled “GLM vs Kimi”| Model | AA Index | SWE-bench Verified | SWE-bench Pro | Input $/M | Output $/M | Context | Weights |
|---|---|---|---|---|---|---|---|
| GLM 5.3 | 45 | — | — | $1.40 | $4.40 | 1M | open |
| Kimi K2.6 | 27 | 80.2 | 58.6 | $0.95 | $4.00 | 256K | open |
| Kimi K2.7 | 26 | — | — | $0.95 | $4.00 | 256K | open |
Both are open heavyweight coders; the split is horizon. GLM 5.3 scores 45 on the AA index — a point above Kimi K3 at under a third of its price — and has a longer context window; reach for it on long-horizon, whole-repo work. Kimi holds the best open SWE-bench Verified score and is tuned for tool-heavy agent runs. Per-token pricing is close, so choose on the work rather than the sticker: GLM for the longest, hardest tasks, Kimi for multi-step agentic coding.
GLM 5.3 is on our pools flat-rate (the Kimi K2 line was retired; Kimi K3 is served in the Flagship Pool) — capacity you reserve, not metered volume.
GLM vs DeepSeek
Section titled “GLM vs DeepSeek”| Model | AA Index | SWE-bench Verified | SWE-bench Pro | Input $/M | Output $/M | Context | Weights |
|---|---|---|---|---|---|---|---|
| GLM 5.3 | 45 | — | — | $1.40 | $4.40 | 1M | open |
| DeepSeek V4.1 Flash | 39 | — | — | $0.30 | $1.20 | 1M | open |
Weight classes again. GLM 5.3 is a heavyweight for the hardest coding and long-horizon reasoning; DeepSeek V4.1 Flash is the budget anchor for volume. GLM costs more per token but scores higher on the index; DeepSeek matches mid-tier closed models on the index at budget-tier pricing. Reach for GLM 5.3 when the task is hard enough that the extra capability pays for itself, DeepSeek when you are running all day and per-token cost compounds.
Both GLM 5.3 and DeepSeek V4.1 Flash are available flat-rate on our pools — no token caps during your reserved hours.
MiniMax vs DeepSeek
Section titled “MiniMax vs DeepSeek”| Model | AA Index | SWE-bench Verified | SWE-bench Pro | Input $/M | Output $/M | Context | Weights |
|---|---|---|---|---|---|---|---|
| MiniMax M3 | 29 | 80.5 | 59.0 | $0.30 | $1.20 | 1M | open |
| DeepSeek V4.1 Flash | 39 | — | — | $0.30 | $1.20 | 1M | open |
Both open, both long-context, but different tiers. MiniMax M3 sits near the top of the open SWE-bench Verified scores at the lowest price in that band and adds native multimodal input — reach for it on whole-codebase, long-document or image and video work. DeepSeek V4.1 Flash is cheaper still and tuned for high-volume pipelines. Choose MiniMax when you need the capability or the modalities, DeepSeek when the job is bulk text and cost is the deciding factor.
MiniMax M3 and DeepSeek V4.1 Flash are both served flat-rate on our pools — you buy reserved capacity, not metered tokens.
Caveats
Section titled “Caveats”Scores are as published by each vendor or independent evaluator; scaffolding differs between labs, which is why every number links to its source. The AA Intelligence Index is one composite — task-specific rankings differ (see the Kimi callout). Prices are list prices for the variants AA evaluates; open models are often cheaper on aggregators. Where a vendor publishes no score, the cell shows — rather than an estimate.
Changelog
Section titled “Changelog”-
2026-09-23 — the budget-multimodal recommendation moves to MiMo V2.6 Flash, which replaced MiMo V2.5 in our Core Pool in place at the same $0.14/$0.28 list price: same size class (309B MoE / 15B active), a generation of RL post-training (vendor agent benchmarks within 4 points of the 1T MiMo-V2.6-Pro), image input, MIT weights. No AA index yet, so both MiMo rows stay index-less in the budget table rather than estimated.
-
2026-09-22 — AA moved its Intelligence Index to v4.3.2: every previously tracked model on both tables fell 9–19 points at unchanged prices (Opus 5 51, Fable 5 50, GPT-5.6 Sol 47, K3 44); GPT-6 Astra and Claude Fable 5.1 (53) join the heavyweight table as the new ceiling — treat pre/post scores as different scales. MiMo-V2.6-Pro joins the heavyweight table as the top open model (46, $0.43/$0.87), 7 points off the ceiling. DeepSeek V4.1 Flash is scored (39) and replaces V4 Flash in the budget table and the head-to-heads; MiMo-V2.6-Flash (under review as MiMo V2.5’s successor) joins the budget table without an index yet.
-
2026-09-10 — the high-volume recommendation moves to DeepSeek V4.1 Flash, the new generation that replaced V4 Flash in our Core Pool: 552B MoE (8B active in prefill, 16B in decode), causal encoder–decoder, native image input, 1M context, MIT weights, at a lower list price than the model it replaces. The budget table still shows V4 Flash’s measured numbers — Artificial Analysis has not scored V4.1 yet, and we don’t estimate.
-
2026-08-30 — GPT-5.6 Sol: $5/$30 idx 61 → $4/$20 idx 61.
-
2026-08-14 — AA recalibrated its Intelligence Index to v4.1.1: every model on both tables shifted up 1–5 points at unchanged prices (Opus 5 63, Fable 5 62, GPT-5.6 Sol 61, Kimi K3 60) — the ordering is essentially preserved, so treat pre/post-recalibration scores as different scales. Qwen3.8 Max joins the heavyweight table (58, $2/$6) now that its weights are open. DeepSeek V4 Flash’s list price rose to its new peak rate ($0.44/$1.32), the one price move of the collection.
-
2026-08-01 — DeepSeek V4 Flash re-scored 40 → 50 on the AA index at unchanged prices, following the V4-Flash-0731 retrain.
-
2026-07-28 — release-week update: Kimi K3 (57), Claude Opus 5 (61) and GPT-5.6 Sol (59) join the heavyweight table; new head-to-head Kimi K3 vs K2.7. The open-vs-closed verdict shifts: the gap to the ceiling is now 4 index points. Also collected: Claude Sonnet 5 price cut $3/$15 → $2/$10; Kimi K2.6 index 43 → 44.
-
2026-07-09 — added a Head-to-head section covering the four most-compared pairs (Kimi vs DeepSeek, GLM vs Kimi, GLM vs DeepSeek, MiniMax vs DeepSeek). Each table reads from the same monthly collection as the rest of the guide, so the numbers move with it; the verdicts are static.
-
2026-07-06 (first edition) — this guide absorbs and supersedes two blog posts (“Open-source vs proprietary LLMs in 2026” and “LLM weight classes”), whose data now lives here and refreshes monthly instead of rotting in dated posts. Baseline verdicts: heavyweight ceiling closed (Opus 4.8 / Fable 5), everything below it open (GLM 5.2, MiniMax M3, Kimi K2.6/K2.7); budget division dominated by open (DeepSeek V4 Flash, MiMo V2.5).
Companion reports: LLM Pareto Frontier (the price-intelligence map) · State of Open Weights (the served models in depth) · LLM Price Tracker (prices over time).