Skip to content

Blog

DeepSeek V4-Flash-0731: what changed, and how to use it over the API

On July 31, DeepSeek shipped V4-Flash-0731 — not a bigger model, a retrained one. Same 284B-total / 13B-active architecture, same 1M-token context, same price bracket — but re-post-trained for agent work, and the reported jump is unusual: DeepSeek says the new Flash now beats its own larger V4-Pro-Preview on every one of the nine agent benchmarks it publishes.

If you use the Core Pool, there is nothing to migrate: the model id is still deepseek-v4-flash, and requests already serve the new build. Unlimited, flat-rate access starts from $14.99/mo ($12.74/mo billed annually) — live pricing on /pools.

Same architecture, new post-training. This is DeepSeek’s own reported before/after (vendor harness — no independent verification of these specific numbers yet):

V4-Flash-0731V4-Flash preview (the previous build, April 2026)
Terminal-Bench 2.1 DSBench-FullStack DeepSWE V4-Flash-0731 — Terminal-Bench 2.1: 82.7 V4-Flash preview — Terminal-Bench 2.1: 61.8 V4-Flash-0731 — DSBench-FullStack: 68.7 V4-Flash preview — DSBench-FullStack: 37.0 V4-Flash-0731 — DeepSWE: 54.4 V4-Flash preview — DeepSWE: 7.3 82.761.8 68.737.0 54.47.3

Full vendor-stated table for the 0731 build, next to the April preview it replaces:

BenchmarkFlash-0731Flash preview (Apr)
Terminal-Bench 2.182.761.8
Cybergym76.738.7
Toolathlon-Verified70.349.7
DSBench-FullStack†68.737.0
DSBench-Hard†59.625.8
DeepSWE54.47.3
NL2Repo54.239.4
Agents’ Last Exam25.215.8
AutomationBench Public25.110.8

† DeepSeek-internal test sets; the rest are public benchmarks.

The DeepSWE number is the striking one: the retrain multiplied the preview’s score by seven without touching the architecture. And like the April preview, the 0731 checkpoint is a genuine open-weights release: the weights are published under MIT at deepseek-ai/DeepSeek-V4-Flash-0731 (304B parameters on the repo — the 284B base plus a draft module).

The independent signal comes from Artificial Analysis, which measures Flash-0731 at 50 on its Intelligence Index — ten points above the previous Flash. Honest framing: it does not beat Claude Opus 5 (61) or GPT-5.6 Sol (59) — nothing near this price does. What it does is land within seven points of Kimi K3 (57, the highest open-weights score on the index) from the budget model of the field — with a reference per-token price of $0.28 per 1M output tokens, roughly 90× below Claude Opus 5’s $25:

Served in a CheapestInference poolReference frontier models
Claude Opus 5 Claude Fable 5 GPT-5.6 Sol Kimi K3 Claude Opus 4.8 GLM 5.2 V4-Flash-0731 MiniMax M3 Claude Opus 5 (max effort): 61 — list output $25.00/1M Claude Fable 5 (max effort, Opus 4.8-fallback config as evaluated by AA): 60 — list output $50.00/1M GPT-5.6 Sol (max): 59 — list output $30.00/1M Kimi K3 (max) — Flagship Pool: 57 — list output $15.00/1M Claude Opus 4.8: 56 — list output $25.00/1M GLM 5.2 (max) — Frontier Pool: 51 — list output $4.40/1M DeepSeek V4-Flash-0731 — Core Pool: 50 — list output $0.28/1M MiniMax M3 — Frontier Pool: 44 — list output $1.20/1M 6160 5957 5651 5044

Same data as a table, with list output prices alongside for scale — each model at its best published configuration (max effort where the index reports one); the median across all models Artificial Analysis tracks is 17:

ModelAA Intelligence IndexList output $/1MOn CheapestInference
Claude Opus 561$25.00
Claude Fable 5*60$50.00
GPT-5.6 Sol59$30.00
Kimi K357$15.00Flagship Pool
Claude Opus 4.856$25.00
GLM 5.251$4.40Frontier Pool
DeepSeek V4-Flash-073150$0.28Core Pool
MiniMax M344$1.20Frontier Pool

* Yes, one point below Claude Opus 5, even though Anthropic positions Fable 5 above Opus in capability. The index scores what AA actually evaluates: Fable 5 is measured in its “Adaptive Reasoning, Max Effort, Opus 4.8 Fallback” serving configuration — the variant with Anthropic’s additional dual-use safeguards — while Opus 5 runs at plain max effort. Independent index, published configurations, taken as-is.

Four of the eight models on that chart are served here on flat-rate subscriptions — and the 0731 retrain just moved the cheapest pool of the three into frontier territory.

V4-Flash-0731 vs Claude Opus 4.8, GPT-5.6 Luna, and Gemini 3.6 Flash

Section titled “V4-Flash-0731 vs Claude Opus 4.8, GPT-5.6 Luna, and Gemini 3.6 Flash”

Is DeepSeek V4-Flash-0731 better than Claude Opus? On raw capability, no — and we won’t pretend otherwise. The strongest competitor in DeepSeek’s own release table is Claude Opus 4.8, and Opus wins every one of the nine shared benchmarks. What the table actually shows is how little it wins by, against a model priced roughly 90× higher per output token at list:

Benchmark (DeepSeek’s release table, vendor-run)V4-Flash-0731Claude Opus 4.8
Terminal-Bench 2.182.785.0
Cybergym76.783.1
Toolathlon-Verified70.376.2
DSBench-FullStack†68.771.6
DSBench-Hard†59.671.7
DeepSWE54.458.0
NL2Repo54.269.7
Agents’ Last Exam25.225.7
AutomationBench Public25.127.2
Reference output price, per 1M tokens$0.28$25.00

On Agents’ Last Exam the gap is half a point — effective parity. On Terminal-Bench 2.1 it is 2.3 points. NL2Repo and the (internal) DSBench-Hard are where the distance stays wide.

Against its actual price peers, the independent picture flips. Per Artificial Analysis: Flash-0731 sits one point behind GPT-5.6 Luna (50 vs 51 at max effort) with a cost per task roughly 60% lower — even after OpenAI’s price cut — ties Gemini 3.6 Flash (50), and lands on AA’s Pareto frontier for Intelligence vs Cost per Task: at this intelligence level, nothing tracked is cheaper per task.

Here is that frontier drawn out — intelligence against list output price. A model is on the frontier when nothing tracked is both smarter and cheaper; everything below-right of the line pays more for less. The 0731 retrain moved Flash onto it, and pushed MiniMax M3 off:

Served in a CheapestInference poolReference frontier modelsPareto frontier
4045 5055 60 $0$10 $20$30 $40$50 List output price — $ per 1M tokens AA Intelligence Index ↑ DeepSeek V4-Flash-0731 — Core Pool: index 50 at $0.28/1M — on the frontier MiniMax M3 — Frontier Pool: index 44 at $1.20/1M GLM 5.2 — Frontier Pool: index 51 at $4.40/1M — on the frontier Gemini 3.5 Flash: index 50 at $9.00/1M Claude Sonnet 5: index 53 at $10.00/1M — on the frontier Kimi K3 — Flagship Pool: index 57 at $15.00/1M — on the frontier Claude Opus 4.8: index 56 at $25.00/1M Claude Opus 5: index 61 at $25.00/1M — on the frontier GPT-5.6 Sol: index 59 at $30.00/1M Claude Fable 5: index 60 at $50.00/1M (AA config: max effort, Opus 4.8 fallback) V4-Flash-0731 MiniMax M3 GLM 5.2 Gemini 3.5 Flash Sonnet 5 Kimi K3 Opus 4.8 Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Flash-0731 matches Gemini-Flash-class intelligence at 32× lower list price

Three of the five models on that frontier — Flash-0731, GLM 5.2 and Kimi K3 — are served here on flat rate. This chart is a snapshot of the frontier’s cheap end; the monthly-updated, full-field version (with cost-per-task data and edition history) lives in our LLM Pareto Frontier report.

And on a flat-rate subscription the per-token column stops mattering altogether: a Core Pool block is the same price whether your agent burns one million tokens or one billion.

curl:

Terminal window
curl https://api.cheapestinference.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek-v4-flash", "messages": [{"role": "user", "content": "Hello"}]}'

OpenAI SDK (Python):

from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="sk-...", # your subscriber key
)
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Fix the failing test in this repo..."}],
)
print(response.choices[0].message.content)

No account yet? Register, subscribe to the Core Pool, and mint a key at cheapestinference.com/keys.

The API also speaks the Anthropic Messages format, so a retrained agent model drops straight into the most popular coding agent:

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-..." # your subscriber key
export ANTHROPIC_MODEL="deepseek-v4-flash"
export ANTHROPIC_SMALL_FAST_MODEL="deepseek-v4-flash"

Start claude as usual — every request runs on Flash-0731 with no per-token meter. The same key works in Cline, Roo Code, Continue, and anything that accepts a custom OpenAI base URL (setup guides).

Why this release matters for flat-rate users

Section titled “Why this release matters for flat-rate users”

Agent workloads are exactly where per-token bills explode: an agent re-sends its growing context on every tool call, and iteration count — not task value — drives the invoice. A model that is suddenly much better at agent work makes that math worse per-token and better flat-rate: more capable loops, same fixed monthly price. We ran the full per-token vs. flat break-even math in Unlimited DeepSeek: what a flat monthly subscription changes — every row of that table just got more favorable, because the same subscription now serves a stronger model.

The trade-offs, as always: usage is unlimited in tokens during your reserved 8-hour blocks, each subscription runs one request at a time (fair use), and outside your blocks the key doesn’t serve. Full details: DeepSeek V4 Flash API docs · Plans & Limits.

Check live Core Pool availability →


CheapestInference serves Kimi K3 (Flagship Pool), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

How to use the Kimi K3 API: key, snippets, and unlimited access

Kimi K3 is live on CheapestInference: Moonshot’s flagship — scoring 57 on the Artificial Analysis Intelligence Index, just 4 points off the top model and the narrowest open-vs-closed gap since February — served through an OpenAI- and Anthropic-compatible API with unlimited usage at a flat monthly price, from $149/mo (live availability — seats are very limited). This is the practical guide: get access, call it, wire it into your coding agent.

Three realistic routes to K3 over an API today:

  1. Per-token, from Moonshot — $3.00 per 1M input tokens ($0.30 cached) and $15.00 per 1M output. Ideal for evaluation and light use; expensive fast for agent workloads.
  2. Kimi memberships — Moonshot’s own plans bundle K3 access with request quotas per 5-hour and weekly windows (new signups have been intermittently paused since launch).
  3. Unlimited time-block subscription (this guide) — reserve one or more daily 8-hour blocks on the Flagship Pool and use K3 with no token caps during your hours. From $149/mo per block; all three blocks = 24/7.

For the subscription route: create an account, subscribe to the Flagship Pool, and mint an API key at cheapestinference.com/keys. Model id: kimi-k3.

curl:

Terminal window
curl https://api.cheapestinference.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "kimi-k3", "messages": [{"role": "user", "content": "Hello"}]}'

OpenAI SDK (Python):

from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="sk-...", # your subscriber key
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Refactor this function..."}],
)
print(response.choices[0].message.content)

K3 is a reasoning model, but reasoning is off by default — you get fast, direct answers. Turn it on per request with "thinking": {"type": "enabled"} or reasoning_effort (low / high / max — K3’s deepest tier). With reasoning on, give it a generous max_tokens so long answers don’t truncate mid-thought, and stream (stream: true) for responsive UIs.

The API also speaks the Anthropic Messages format, so Claude Code works out of the box:

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-..." # your subscriber key
export ANTHROPIC_MODEL="kimi-k3"
export ANTHROPIC_SMALL_FAST_MODEL="kimi-k3"

Start claude as usual — every request now runs on K3 with no per-token meter. The same key works in Cline, Roo Code, Continue, and any client that accepts a custom OpenAI base URL (setup guides).

At list price, a coding agent that burns 100M tokens a month on K3 costs roughly $200–400 per-token, depending on your cache-hit and output mix — and heavy agentic users go far beyond that. A Flagship block at $149/mo covers your working day; all three blocks cover 24/7. The trade-offs: during your reserved hours usage is truly unlimited, each subscription runs one request at a time (fair use), and outside your blocks the key doesn’t serve. Full pricing detail: Plans & Limits · Kimi K3 API docs.

Seats on the Flagship Pool are very limited — when a block sells out it’s gone until someone leaves. Check live availability →


CheapestInference serves Kimi K3 (Flagship Pool), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Kimi K3: specs, benchmarks, and our day-one plan to serve it

Kimi K3 is Moonshot AI’s new flagship, announced on July 16 after a leaked promotion page on Moonshot’s own platform tipped the release a day early. The headline is simple: on the independent Artificial Analysis Intelligence Index it scores 57 — above Claude Opus 4.8 (56), making it the first open-weight model to outscore a Claude Opus-class model on that index.

And to answer the question this blog exists to answer: yes, we will serve it. We have already secured the capacity to run a model of this size. K3 goes live on CheapestInference the moment two things are true: the weights are actually published, and the license terms permit commercial serving. Nothing else is in the way.

Update — July 2026: Kimi K3 is now live. The weights shipped on schedule (July 27), and K3 is served in the new Flagship Pool with unlimited usage at a flat monthly price — very limited seats. How to use the Kimi K3 API → · Subscribe →

July 28: with Claude Opus 5 (61) and GPT-5.6 Sol (59) landing the same month, K3’s 57 puts the open-vs-closed gap at 4 index points — per Artificial Analysis, the narrowest since the GLM-5 release in February. Only two labs score higher. Our Pareto Frontier report now charts K3 — straight onto the frontier.

ArchitectureMixture-of-Experts, ~2.8T total parameters
Context window1M tokens
InputText, image, and video
Variants at launchK3 Max (chat and agent tasks) · K3 Swarm Max (large-scale parallel processing)
Available todayMoonshot’s API, Kimi Code, and the Kimi app
Open weightsPublished July 27, 2026Hugging Face, Kimi K3 License
List price (API)$3 input / $15 output per 1M tokens

Coverage from the launch day: TechCrunch on the Opus gap closing, Fortune on Chinese AI entering Fable-level territory, and Simon Willison’s notes for a practitioner’s first look.

K3 was announced on July 16, 2026, usable from day one through Moonshot’s own API, Kimi Code, and the Kimi app. The open weights were published on July 27, 2026, on schedule, on Hugging Face under Moonshot’s own Kimi K3 License — and K3 went live on CheapestInference’s Flagship Pool the same week.

  • Artificial Analysis Intelligence Index: 57. For scale: Claude Opus 4.8 scores 56, and the best open-weight model until now — GLM 5.2 — scores 51: the open ceiling jumped six points in one release. Only three closed models score higher — Claude Opus 5 (61), Claude Fable 5 (60) and GPT-5.6 Sol (59). Where every model sits on price-vs-intelligence is our Pareto Frontier report, whose refreshed 2026-07 edition now charts K3 — straight onto the Pareto frontier.
  • #1 on Frontend Code Arena with 1,679 points — ahead of Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and GLM 5.2 (1,587).
  • The Kimi track record. Moonshot’s K2.6 still holds the best open SWE-bench Verified score (80.2), and the K2 line has been the default open choice for tool-heavy agent work — see the tier-fair matchups in our Which-LLM guide.

The honest caveats, same as in our living reports: the index is one composite and task-specific rankings differ, and these are launch-week numbers, mostly on Moonshot’s own serving stack. The weights question resolved on schedule: K3 is downloadable from Hugging Face under Moonshot’s own Kimi K3 License (not the K2 line’s Modified MIT — read the text before self-hosting), and K3 now has its row in State of Open Weights.

When K3 was announced, every model above GLM 5.2’s intelligence score was closed, and the open-vs-closed gap had sat at 5+ index points all year. With the weights shipped, the open ceiling is 57 — and after Claude Opus 5 (61) and GPT-5.6 Sol (59) landed in the same month, the gap to the very top stands at 4 index points, the narrowest since the GLM-5 release in February (Artificial Analysis). K3’s $3/$15 list price still undercuts every closed model in its class, and it walked straight onto the Pareto frontier.

Moonshot’s list price for the K3 API is $3 input / $15 output per 1M tokens — undercutting every closed model in its class, but 3–4× the K2 line’s price: the first open flagship priced like a closed mid-tier model (Price Tracker). On CheapestInference that call is made: K3 debuted in its own Flagship Pool, flat-rate from $149/mo, unlimited usage during your reserved hours and very limited seatsthe pools page is always the live source for lineup, prices and availability.

  • Live now. The capacity we secured before launch is serving K3 today in the Flagship Pool — we didn’t start the clock the day the weights dropped.
  • The usual: one OpenAI- and Anthropic-compatible API, flat-rate time-block subscriptions, no token caps during your reserved hours — so it drops into Claude Code, Cline, or any compatible client; it’s in GET /v1/models now.

If you want to be running K3 this week, create an account — seats in the Flagship Pool are very limited, no waitlist.


CheapestInference serves Kimi K3 (Flagship Pool, from $149/mo), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool, from $48.45/mo billed annually) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool, from $12.74/mo) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Combined keys: stack your subscriptions into one credential

On an Unlimited subscription, the thing you’re actually shaping isn’t tokens — it’s capacity over time. Each subscription gives you a daily coverage window (the 8-hour blocks you reserve) and, within it, some number of requests you can run in parallel. Your real capacity is those two dimensions multiplied: parallel slots × hours.

Most people grow along both axes. You buy a second block to cover more of the day, or a second subscription to run more requests at once during your busy hours. The awkward part used to be the bookkeeping: every subscription minted its own key, so a serious setup meant three or four keys, each live at different times, each with its own capacity — and you juggling which one to paste where.

Combined keys remove the juggling. A combined key (it looks like sk-ci-meta-…) folds two or more of your subscriptions into a single credential, and it behaves like the union of everything underneath it.

Coverage adds up. A combined key is live whenever any of its subscriptions has an open window. Reserve the Europe block on one subscription and the Americas block on another, combine them, and the one key covers both stretches of the day.

Overlap stacks parallel capacity. Where two subscriptions cover the same hour, their parallel slots add together. That’s the lever for concurrency: if you need to run more requests side by side during your peak hours, buy a second subscription over those hours and combine it in.

A concrete example. Say you hold one full-day (24h) subscription and add a second subscription on just the Europe block. The combined key gives you:

  • Double the parallel capacity during the Europe hours — the two subscriptions overlap there, so their slots stack.
  • Baseline capacity the rest of the day — only the 24h subscription is covering those hours.

Coverage is 24/7 (from the full-day subscription); parallel capacity is shaped to peak exactly when you work.

Combining is about capacity, not billing. Each subscription keeps its own monthly allowance and its own renewal date — nothing is pooled or co-mingled. The practical upside is resilience: if one subscription lapses or you let it cancel, it simply drops out of the union. The combined key keeps working with whatever subscriptions remain — no dead key, no scramble to re-issue credentials.

Requests route by the model you ask for. So a combined key can span subscriptions on different pools — a Core Pool subscription and a Frontier Pool subscription under one key — and each request lands wherever its model lives. Ask for deepseek-v4-flash and it’s served from your Core subscription; ask for kimi-k2.7 and it’s served from your Frontier one. GET /v1/models on a combined key lists every model you can reach across all of them.

Combining a subscription into a key removes that subscription’s standalone key — each subscription has exactly one credential at a time. That’s deliberate: it means there’s never ambiguity about what capacity a given key carries. A key’s coverage and parallel capacity are always the exact sum of the subscriptions it holds, nothing more, nothing hidden. The dashboard walks you through it when you combine.

In the Keys page, choose Create API Key and multi-select the subscriptions you want to combine. Before you commit, a preview shows the resulting daily coverage and peak parallel capacity, so you can see the shape you’re buying into. Prefer the API? The Management API does the same thing programmatically.

Full walkthrough in the combined keys guide. If you’re still deciding which blocks to reserve, the live menu and per-block prices are on the pools page.


CheapestInference serves frontier open-weights models — Kimi K2.7, GLM 5.2, MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash, MiMo v2.5 (Core Pool) — through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Qwen coding plans in 2026: what you actually get

Qwen3-Coder — Alibaba’s open-weights coder family — is one of the most capable coding models you can run over an API, and one of the most searched-for. If you want to run it as your daily coding model, you have a few realistic routes. As with Kimi and GLM, they differ less in headline price than in cost shape: what happens to your bill and your workflow when a heavy week hits.

Evaluating Qwen coding plans? Here’s how flat-rate unlimited access to comparable open-weights models stacks up. We don’t serve Qwen — but if what you’re really after is a fixed monthly bill for a capable open-weights coder, the same cost-shape question applies to models like Kimi K2.7, GLM 5.2, DeepSeek V4 Flash, MiMo v2.5, and MiniMax M3. Compare the flat-rate pools →

Alibaba Cloud sells a subscription Coding Plan on Model Studio aimed specifically at coding-tool usage of Qwen (plus a few third-party models). It’s first-party access, integrated with Qwen Code and compatible with Claude Code, Cline, and Cursor, and the Qwen coder models — qwen3-coder-plus, qwen3-coder-next, and the newer qwen3.x-plus line — arrive there first.

The trade-off is that the plan is quota-based: each tier grants a request allowance that resets on a schedule — with per-few-hours, weekly, and monthly request caps — and burning through it mid-refactor means waiting for the reset or moving up a tier. The tier lineup itself has already shifted once in 2026 (the entry-level Lite tier stopped accepting new orders), so any number printed here would go stale — check Alibaba’s Coding Plan page for the current tiers and quotas.

Good fit: you want first-party access to the newest Qwen coder models and your volume fits inside a tier’s request quota.

Qwen3-Coder is available per-token from Alibaba’s Model Studio / DashScope API and from several aggregators. No tiers, no resets — you pay for exactly the tokens you burn, which is ideal while you’re evaluating the model or your usage is light. New Model Studio accounts also get a time-limited free-token trial (region-restricted, expiring after a fixed window), useful for a first look — see Alibaba’s pricing page for the current allowance.

The catch is structural, not Qwen-specific: coding agents re-send their whole context on every tool call, so token volume compounds with every iteration. A capable coder will happily churn through long agent sessions — great for output, open-ended for the invoice. Per-token Qwen is cheap per request and unpredictable per month.

Good fit: a few million tokens a month, spiky schedules, or benchmarking before committing.

Route 3: unlimited time blocks on comparable open-weights models

Section titled “Route 3: unlimited time blocks on comparable open-weights models”

The third shape is the one we sell, so apply the usual discount for self-interest — and note the honest caveat up front: we don’t serve Qwen. What we offer is the same cost shape for a lineup of comparable open-weights coders. You reserve one or more daily 8-hour time blocks and get unlimited usage during them: no token allowances, no request quotas, no resets — a monthly number that’s fixed the day you subscribe. Capacity is shaped by per-key concurrency instead of token or request budgets, so an agent that loops all afternoon changes nothing on the bill.

Two things matter if you’re weighing this against a Qwen plan:

  1. Comparable models, one fee. A Frontier Pool block covers Kimi K2.7, GLM 5.2, and MiniMax M3 (1M context); the Core Pool covers DeepSeek V4 Flash and MiMo v2.5 — all open-weights, switchable per request. If you were choosing Qwen for capability-per-dollar, these are in the same class.
  2. It runs in the same tools. The API speaks both the Anthropic and OpenAI formats, so it drops into Claude Code, Cline, Roo Code, or Qwen Code — no wrapper, just a base-URL change.

Current block pricing is on the pools page.

Good fit: you want a fixed monthly bill for a capable open-weights coder and predictable working hours — and you’re not locked to the Qwen name specifically.

RouteCost shapeLimitsModelsBest for
Official Qwen Coding PlanFixed monthly subscriptionRequest quotas that reset (per-few-hours / weekly / monthly)First-party Qwen coder models (plus some third-party)First-party Qwen, volume inside quota
Per-token APIPay per token usedNone — spend scales with usageAny Qwen model on Model Studio / aggregatorsLight, spiky, or exploratory use
Flat-rate time blocks (us)Fixed monthly, per blockConcurrency-shaped; no token or request capsComparable open-weights (Kimi, GLM, DeepSeek, MiMo, MiniMax) — not QwenPredictable hours, fixed bill, model-agnostic

Official Qwen Coding Plan — first-party access, day-one Qwen coder updates, your volume fits the request quota. Per-token — light, spiky, or exploratory usage; pay only for what you burn. Flat-rate time blocks — heavy daily coding in predictable hours where you want a constant bill and you’re open to a comparable open-weights model rather than Qwen specifically.

All three answer the same underlying question. It isn’t “which Qwen tier is cheapest” — it’s which cost shape matches how you work, and whether you need the Qwen name or just a capable open-weights coder at a fixed price.


CheapestInference serves Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. We do not serve Qwen. See the pools or get started.