Skip to content

Blog

MiMo-V2.6 Pro & Flash: specs, benchmarks, pricing & API options

Update — September 23, 2026: MiMo V2.6 Flash is now served in the Core Pool as mimo-v2.6-flash, replacing MiMo v2.5 in place — same pool, same price, subscriptions uninterrupted; the old id mimo-v2.5 keeps working until October 23, 2026. Model page: MiMo V2.6 Flash API. MiMo-V2.6-Pro remains under evaluation (live status).

Xiaomi’s MiMo team shipped two models on September 21, 2026, and each one matters for a different reason:

  • MiMo-V2.6-Pro is the new #1 open-weights model on the Artificial Analysis Intelligence Index: 46, ahead of GLM-5.3 and Qwen3.8 Max (45) and Kimi K3 (44), at a blended $0.18 per 1M tokens.
  • MiMo-V2.6-Flash is the direct successor to MiMo v2.5: the same size class (309B total / 15B active) and the same list price, with agent-benchmark scores that land within a few points of Pro.

Both ship MIT-licensed weights, a 1M-token context window and omnimodal input (text, image, audio, video). A naming note for searchers: this is MiMo-V2.6, September 2026. It is not the older MiMo-V2-Flash from late 2025, which had a 256K context and is a different model.

MiMo-V2.6-ProMiMo-V2.6-Flash
Architecture1.02T total / 42B active MoE309B total / 15B active MoE, hybrid sliding-window + global attention
DecodingReasoning modelReasoning model, 5-layer multi-token prediction (7 tokens per pass)
InputText, image, audio, videoText, image, audio, video
Context window1M tokens, up to 128K output1M tokens, up to 128K output
Open weightsXiaomiMiMo/MiMo-V2.6-Pro-RL, MIT, FP8XiaomiMiMo/MiMo-V2.6-Flash-RL, MIT, BF16 + FP8
List price (API)$0.435 in / $0.87 out per 1M, cached input $0.0036$0.14 in / $0.28 out per 1M, cached input $0.0028
Measured speed124.5 output tok/s, 2.34s to first token (AA)Not yet measured independently

The “-RL” suffix in the repository names confuses people. These are full checkpoints from the final reinforcement-learning run, not adapters. Pro and Flash are also separate training runs with their own benchmark tables, not one model cut down to two sizes.

These are Xiaomi’s own reported numbers from its vendor harness, not independently reproduced. The baseline is last generation’s largest model, MiMo-V2.5-Pro:

MiMo-V2.5-Pro (previous generation)MiMo-V2.6-Flash (15B active)MiMo-V2.6-Pro (42B active)
Terminal-Bench 2.1 DeepSWE v1.1 CyberGym MiMo-V2.5-Pro — Terminal-Bench 2.1: 65.2 MiMo-V2.6-Flash — Terminal-Bench 2.1: 87.6 MiMo-V2.6-Pro — Terminal-Bench 2.1: 89.9 MiMo-V2.5-Pro — DeepSWE v1.1: 19.0 MiMo-V2.6-Flash — DeepSWE v1.1: 67.9 MiMo-V2.6-Pro — DeepSWE v1.1: 71.9 MiMo-V2.5-Pro — CyberGym: 40.0 MiMo-V2.6-Flash — CyberGym: 95.1 MiMo-V2.6-Pro — CyberGym: 94.0 65.287.689.9 19.067.971.9 40.095.194.0
Benchmark (Xiaomi, vendor-run)MiMo-V2.5-ProMiMo-V2.6-FlashMiMo-V2.6-Pro
Terminal-Bench 2.165.287.689.9
DeepSWE v1.119.067.971.9
CyberGym40.095.194.0
AutomationBench16.052.353.1
MiMo Code Bench40.461.263.2
JobBench—61.262.0
MiMo Visual Coding—71.572.3

The gap between the two new models is small. Flash is within 4.0 points of Pro on every benchmark, and it beats Pro on CyberGym. The generational jump is large: on DeepSWE, Flash scores 67.9 where last generation’s Pro scored 19.0. Xiaomi credits scaled RL for this: about 750,000 trajectories across more than 7,000 environments, and it published the runs live while they trained. That jump is the headline, and it is also the reason to wait for independent reproductions before trusting every point.

The independent signal comes from Artificial Analysis, using Intelligence Index v4.3.2. Every score below was checked today on that edition. AA rescored the whole field in this edition, so these numbers are not comparable with scores quoted from earlier editions.

MiMo-V2.6-ProServed in a CheapestInference poolClosed frontier reference
GPT-6 Astra · Fable 5.1 MiMo-V2.6-Pro GLM-5.3 Qwen3.8 Max Kimi K3 DeepSeek V4.1 Flash GPT-6 Astra: 53 — top of the index, tied with Claude Fable 5.1 MiMo-V2.6-Pro: 46 — #1 open-weights model, blended $0.18/1M (AA) GLM-5.3 (max) — Frontier Pool: 45, blended $0.90/1M (AA) Qwen3.8 Max — Flagship Pool: 45, blended $1.18/1M (AA) Kimi K3 (max) — Flagship Pool: 44, blended $2.31/1M (AA) DeepSeek V4.1 Flash — Core Pool: 39, blended $0.18/1M (AA) 5346 4545 4439
ModelAA Intelligence Index v4.3.2AA blended $/1MAA output tok/s
GPT-6 Astra / Claude Fable 5.153——
Claude Opus 551—56.4
MiMo-V2.6-Pro46$0.18124.5
GLM-5.3 (max)45$0.9061.0
Qwen3.8 Max45$1.1839.2
Kimi K3 (max)44$2.3140.6
DeepSeek V4.1 Flash39$0.18225.6

Blended = AA’s 7:2:1 cache-hit/input/output mix. The full field and its history are in our monthly LLM Pareto Frontier report.

The honest read. Pro does not beat the closed frontier: GPT-6 Astra and Claude Fable 5.1 lead at 53, 7 points ahead. Its lead over the next open-weights models is a single point, which is within the noise of any index. The notable part is the rest of the row. Against the three open-weights models it edges out:

  • it is 5–13× cheaper on AA’s blended price;
  • it is about 2–3× faster on measured output speed.

It costs the same blended $0.18 as DeepSeek V4.1 Flash while scoring 7 points higher. That combination puts Pro on AA’s intelligence-vs-cost Pareto frontier at $0.13 per Index task.

MiMo-V2.6-Flash: a new generation at the old price

Section titled “MiMo-V2.6-Flash: a new generation at the old price”

Flash has no AA score yet. Here is why it may be the more important release for high-volume workloads:

  • The cost profile is unchanged. Flash keeps the exact list price of MiMo v2.5 ($0.14 in / $0.28 out, cached input $0.0028) and the same 15B-active footprint. Agent loops, sub-agents and long-context summarization get a generation’s worth of capability at the same per-token cost.
  • It lands close to Pro. On Xiaomi’s own table, Flash sits within 4.0 points of the 1T Pro everywhere. At the size class, this is the part to check once independent numbers land.
  • 1M context stays usable. The hybrid sliding-window + global attention keeps long inputs affordable to serve, and the 7-token multi-token-prediction head is built for fast decoding.

Pro and Flash both ship under plain MIT on Hugging Face. Pro is in FP8; Flash is in BF16 and FP8. Alongside them, Xiaomi published:

  • a technical report;
  • more than 7,000 RL environments;
  • its end-to-end RL framework;
  • a 9B distilled model.

The full task datasets have not been released yet. For anyone evaluating models to serve rather than just call, the top of the open-weights index now carries the most permissive license there is.

MiMo-V2.6-Flash is served. Since September 23, 2026 it is the Core Pool’s MiMo model, as mimo-v2.6-flash, next to DeepSeek V4.1 Flash, unlimited on flat monthly time blocks. It replaced MiMo v2.5 in place: same pool, same price, nothing to change on existing keys, and requests with the old id mimo-v2.5 are served by V2.6 Flash until October 23, 2026. Image input works on the same endpoint; reasoning is off by default and switched on per request. MiMo-V2.6-Pro is under evaluation — its live status is on the models-under-review page, and any addition lands in the changelog the day it ships.

What is MiMo-V2.6? Xiaomi’s September 2026 model generation, in two open-weights models. MiMo-V2.6-Pro is a 1.02T-total / 42B-active MoE. MiMo-V2.6-Flash is a 309B-total / 15B-active MoE. Both have a 1M-token context window, omnimodal input (text, image, audio, video) and MIT-licensed weights. Xiaomi also offers a Pro-UltraSpeed variant on its own API, advertised at up to 20× Pro’s output speed.

What is the difference between MiMo-V2.6-Pro and MiMo-V2.6-Flash? Size and price. Pro is about 3× larger in active parameters and costs about 3× more per token ($0.435 / $0.87 vs $0.14 / $0.28 per 1M). On Xiaomi’s vendor benchmarks, Flash sits within 4.0 points of Pro on every task and beats it on CyberGym (95.1 vs 94.0).

Is MiMo-V2.6-Pro the best open-weights model? On the Artificial Analysis Intelligence Index v4.3.2 it is #1 among open-weights models, at 46. That is one point ahead of GLM-5.3 and Qwen3.8 Max (45) and two ahead of Kimi K3 (44). It is also the cheapest and fastest of that group: $0.18/1M blended and 124.5 tok/s. GPT-6 Astra and Claude Fable 5.1 lead the overall index at 53.

How much does the MiMo-V2.6 API cost? Xiaomi lists MiMo-V2.6-Pro at $0.435 input / $0.87 output per 1M tokens, with cached input at $0.0036. MiMo-V2.6-Flash is $0.14 / $0.28, with cached input at $0.0028, which is the same as MiMo v2.5.

Does MiMo-V2.6 have open weights? Yes. Both are on Hugging Face under the MIT license: XiaomiMiMo/MiMo-V2.6-Pro-RL and XiaomiMiMo/MiMo-V2.6-Flash-RL. The “-RL” suffix marks the final RL checkpoint; these are full models, not adapters.

Is MiMo-V2.6-Flash the same as MiMo-V2-Flash? No. MiMo-V2-Flash is Xiaomi’s late-2025 model with a 256K context. MiMo-V2.6-Flash is the September 2026 successor to MiMo v2.5, with a 1M context and the RL-scaled post-training described above.

Is there an unlimited MiMo-V2.6 API subscription? Yes, for Flash. Since September 23, 2026, MiMo V2.6 Flash is served unlimited in the Core Pool as mimo-v2.6-flash: flat monthly fee, no token caps during your reserved hours, OpenAI- and Anthropic-compatible, image input included. It replaced MiMo v2.5 in place, and the old id mimo-v2.5 keeps working until October 23, 2026. MiMo-V2.6-Pro is under evaluation (live status) and is not served today.


CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.3 and MiniMax M3 (Frontier Pool) and DeepSeek V4.1 Flash and MiMo V2.6 Flash (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Asking a model "who are you?" proves nothing. Here is what the evidence says.

Every week someone sends a provider the same two screenshots: the request asked for model A, the reply says “I am model B”. The conclusion feels obvious, and it is usually wrong. A subscriber sent us this pair this week, verbatim:

// request
{ "model": "deepseek-v4.1-flash",
"messages": [{ "role": "user",
"content": "Who are you? Who made you? What's your model name, give me a precise model name" }] }
// response
{ "model": "deepseek-v4.1-flash",
"choices": [{ "message": { "role": "assistant",
"content": "I'm DeepSeek, an AI assistant created by the Chinese company DeepSeek (深度求索).\n\nMy precise model name is **DeepSeek-V3**." } }] }

Same subscriber, other runs: the model called itself an Anthropic model. Same key, same endpoint, same served model every time. The model was not lying and the platform was not switching anything. Language models simply do not know what they are, and the research on this is now large enough to put numbers on it.

Study / reportDateWhat was testedResult
TechCrunch: “Why DeepSeek’s new AI model thinks it’s ChatGPT”Dec 2024DeepSeek V3, asked what it isClaimed to be ChatGPT (GPT-4) in 5 of 8 generations; gave OpenAI API instructions when asked about DeepSeek’s API
”I’m Spartacus, No, I’m Spartacus” (Shandong Univ., Drexel, UMass Lowell)Nov 202427 LLMs, systematic identity probing25.93 % exhibit identity confusion; output-distribution analysis attributes it to hallucination, not copied weights
Same paper, documented real-world cases2024Production modelsGemini-Pro says it is Baidu’s Wenxin when asked in Chinese; GPT-4 over the API says it is GPT-3; ByteDance Seed says it is GPT
”Know Thyself? On the Incapability and Implications of AI Self-Recognition” (Univ. of Chicago)Oct 202510 frontier models (GPT-4.1, GPT-5, Claude Sonnet 4, Gemini 2.5 Flash, Kimi K2, DeepSeek V3, GLM-4.5, Qwen3-235B, Grok 4), 20,000 predictionsExact self-identification accuracy 10.3 – 10.9 % against a 10 % random baseline. GPT and Claude receive 97.7 % of all attributions while producing 40 % of the text. GLM-4.5 identified itself as Claude in half of its runs

The Chicago study is the cleanest one: each model was shown text and asked which model wrote it, including its own text. Only four of the ten ever named themselves. The models that were most often named as authors were the ones with the most public visibility, not the ones that actually wrote the sample.

Identity lives in the system prompt, not in the weights. When a vendor’s own chat app says “I am X, made by Y”, that sentence comes from instructions the vendor injects in front of every conversation. Over a raw API, with no system prompt, the model falls back on whatever its training data says about “AI assistants”, and that corpus is dominated by a few famous names.

Training corpora contain other models’ output. The DeepSeek V3 case was traced to exactly this: public datasets full of GPT-4 generated text. Mike Cook, research fellow at King’s College London, described the effect to TechCrunch as “taking a photocopy of a photocopy”. Synthetic data and distillation are standard practice across the industry, so this is not a property of one vendor; the Chicago paper found the same bias in every family it tested.

Version numbers are the least reliable part. A model trained before its own release cannot know its final product name. The Chicago authors report GPT-5 dismissing “gpt-5” as a fake model name because, from the model’s point of view, it had not been released yet. A DeepSeek model calling itself “V3” is the same failure: it names the latest version it read about.

The model field is written by the server. The name in the JSON response is metadata attached by the API, not something the model produced. It tells you what the endpoint claims to have served. It cannot be cross-checked by asking the model, because the model never sees it.

Signals that are more reliable than the model’s own answer

Section titled “Signals that are more reliable than the model’s own answer”

None of these is proof on its own, and each takes some work. All of them beat a “who are you?” prompt, because they measure behaviour the model cannot talk its way around:

  • Tokenizer behaviour. Every model family counts tokens differently, and the tokenizers of open-weights models are public. The usage.prompt_tokens reported for a fixed input is a fingerprint of the tokenizer that actually processed it.
  • Knowledge and capability boundaries. Training cutoffs, supported languages, context length behaviour and native modalities (vision, audio) differ between models in ways that a system prompt cannot fake.
  • Behaviour with structured features. How a model handles tool calls, JSON mode, reasoning fields, stop sequences or unusual parameters is specific to the model and to the software serving it.
  • Response metadata patterns. ID formats, error message wording, header sets and latency profiles tend to be stable per serving stack and change when the stack changes.
  • Consistency across many samples. A single answer at default temperature is noise. The studies above used hundreds or thousands of samples per model; a verification worth trusting does the same and looks at the distribution.
  • Independent, reproducible attestations. Published test harnesses with recorded runs let a third party repeat the measurement. That is how the September 2026 CrofAI investigation established that a reseller was silently routing sixteen model ids to four cheaper models: tool signatures, response ID formats, token counting and canary strings, recorded in public CI runs, and not one “who are you?” prompt.

If a model tells you it is ChatGPT, Claude, or an older version of itself, you have learned something about its training data and nothing about what is being served to you. Treat self-identification the way the research does: as a hallucination category with a measured rate of roughly one model in four, and at chance level when the question gets specific.

When it matters, measure behaviour instead. The signals above are the ones the published investigations relied on, and the ones we would point anyone to, including for our own endpoints.

DeepSeek V4.1 Flash: what changed, and how to use it over the API

On September 10, 2026 DeepSeek released DeepSeek V4.1 Flash — not a retrain this time, a new generation: 552B parameters (mixture-of-experts, 8B active per token in prefill and 16B in decode), a new “causal encoder–decoder” architecture, vision trained in from the start of pre-training, a 1M-token context, and MIT weights on Hugging Face at deepseek-ai/DeepSeek-V4.1-Flash (the tech report PDF ships in the same repo).

It replaced the previous build in our Core Pool the same day. There is exactly one thing to migrate: the model id. The new id is deepseek-v4.1-flash; the old deepseek-v4-flash keeps working as an alias until October 10, 2026, after which it returns an invalid-model error. Same key, same endpoints, same subscription, same reserved hours — unlimited, flat-rate access starts from $22/mo ($18.70/mo billed annually), live pricing on /pools.

Terminal window
# before # from 2026-09-10
"model": "deepseek-v4-flash" "model": "deepseek-v4.1-flash"

The headline claim in DeepSeek’s own release material is that V4.1 Flash beats its own larger V4-Pro on performance, cost, speed and total time to finish a task. It is confident enough in that to act on it: from September 14, 2026 DeepSeek will route its own deepseek-v4-pro traffic to V4.1 Flash, billed at Flash rates, until a V4.1-Pro exists.

Here are the three agent benchmarks from the instruct table on the model card, next to V4-Pro and Claude Opus 5.0. All of these are vendor-run — DeepSeek’s own harness at maximum reasoning effort (reasoning_effort=100), with no independent verification yet:

DeepSeek V4.1-FlashDeepSeek V4-ProClaude Opus 5.0
Terminal-Bench 2.1 DeepSWE v1.1 AutomationBench V4.1-Flash — Terminal-Bench 2.1: 90.6 V4-Pro — Terminal-Bench 2.1: 87.9 Claude Opus 5.0 — Terminal-Bench 2.1: 89.1 V4.1-Flash — DeepSWE v1.1 resolved: 74.2 V4-Pro — DeepSWE v1.1 resolved: 62.7 Claude Opus 5.0 — DeepSWE v1.1 resolved: 74.0 V4.1-Flash — AutomationBench pass@1: 54.8 V4-Pro — AutomationBench pass@1: 43.2 Claude Opus 5.0 — AutomationBench pass@1: 50.3 90.687.989.1 74.262.774.0 54.843.250.3

Same numbers as a table, plus the base-model scores DeepSeek publishes alongside them:

Instruct (vendor harness, max reasoning effort)V4.1-FlashV4-ProClaude Opus 5.0
Terminal-Bench 2.190.687.989.1
DeepSWE v1.1 (resolved)74.262.774.0
AutomationBench (pass@1)54.843.250.3
Base modelV4.1-Flash-BaseV4-Flash-BaseV4-Pro-Base
MMLU-Pro74.168.373.5
HumanEval79.469.576.8
GSM8K93.090.892.6

Two honest readings, and both are worth holding at once. The generous one: on DeepSeek’s own harness, a model with 8–16B active parameters edges past Claude Opus 5.0 on all three agent benchmarks, and past its own Pro-tier sibling on every row of both tables — which is exactly why DeepSeek is willing to point Pro traffic at it. The sceptical one: every number above was produced by the vendor, at maximum reasoning effort, on its own scaffolding. Vendor tables set expectations; they don’t settle them. Where independent evaluation lands is the next section.

Update — September 22, 2026: Artificial Analysis has now scored V4.1 Flash: 39 on its Intelligence Index v4.3.2 — five points above the V4 Flash build it replaced (34 on the same scale), at a lower price. Our Pareto Frontier, Price Tracker and Which-LLM reports now plot it. The paragraph below is the original launch-day note.

Artificial Analysis has not published an Intelligence Index score for V4.1 Flash yet. So we are not moving anything on the strength of a vendor table: our Pareto Frontier, Price Tracker and Which-LLM reports keep the previous build’s plotted point until an independent score exists. For reference, the build this one replaces scored 50 on that index when it launched (52 after AA’s later v4.1.1 recalibration). We’ll update the reports when AA publishes.

The architecture, the vision, and the price

Section titled “The architecture, the vision, and the price”
DeepSeek V4.1 Flash
Parameters552B mixture-of-experts — 8B active per token in prefill, 16B in decode
ArchitectureNew causal encoder–decoder: a 20-layer causal encoder followed by a 20-layer decoder
VisionNative — a from-scratch vision encoder plus projector, trained alongside text from the start of pre-training, not bolted on afterwards
Context window1M tokens
KV cache~890 bytes per token — DeepSeek reports roughly ¼ of the previous generation’s HBM footprint and ⅛ of its SSD footprint
WeightsMIT — deepseek-ai/DeepSeek-V4.1-Flash, tech report in the repo
Model id heredeepseek-v4.1-flash (old deepseek-v4-flash accepted until 2026-10-10)
PoolCore Pool, alongside MiMo V2.6 Flash

The cache line is the one with second-order consequences. A quarter of the KV footprint per token is what makes a 1M-token context economically serious rather than a spec-sheet number — and it is the mechanism behind DeepSeek’s list price coming down on a newer, larger model, which is not the usual direction:

DeepSeek list price, per 1M tokensV4 Flash (before)V4.1 Flash (now)
Off-peak input$0.22$0.15
Off-peak output$0.66$0.60
Peak input$0.44$0.30
Peak output$1.32$1.20
Cache hit (off-peak / peak)—$0.003 / $0.006

Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday; everything else is off-peak. DeepSeek has retired the deepseek-v4-flash id on its own API, where it now routes to V4.1.

On a flat-rate subscription none of that column matters at all — a Core Pool block costs the same whether your agent burns one million tokens or one billion — but it matters as a signal: the cheap tier keeps getting better without getting more expensive. We track that trend month by month in the LLM Price Tracker.

curl:

Terminal window
curl https://api.cheapestinference.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Hello"}]}'

OpenAI SDK (Python):

from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="sk-...", # your subscriber key
)
response = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Fix the failing test in this repo..."}],
)
print(response.choices[0].message.content)

Anthropic SDK (Python) — the same key against the /anthropic endpoint:

from anthropic import Anthropic
client = Anthropic(
base_url="https://api.cheapestinference.com/anthropic",
api_key="sk-...", # your subscriber key
)
message = client.messages.create(
model="deepseek-v4.1-flash",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain this stack trace..."}],
)
print(message.content[0].text)

No account yet? Register, subscribe to the Core Pool, and mint a key at cheapestinference.com/keys.

Claude Code speaks the Anthropic Messages API, so it drops in with two environment variables (plus the small-model one):

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-..." # your subscriber key
export ANTHROPIC_MODEL="deepseek-v4.1-flash"
export ANTHROPIC_SMALL_FAST_MODEL="deepseek-v4.1-flash"

Start claude as usual — every request runs on V4.1 Flash with no per-token meter. Per-project pinning, thinking blocks and the rest of the setup are in the Claude Code + DeepSeek guide. The same key works in Cline, Roo Code, Continue and anything that accepts a custom OpenAI base URL.

Vision is native, and it uses the standard content formats on both endpoints — image_url parts on /v1/chat/completions, image blocks on /anthropic/v1/messages, in user messages, inside the same 8 MB per-request budget as everything else:

response = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What does this dashboard screenshot show?"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
],
}],
)

Does DeepSeek V4.1 Flash have open weights? Yes — MIT, on Hugging Face at deepseek-ai/DeepSeek-V4.1-Flash, with the tech report PDF in the same repo. MIT means unrestricted commercial use, including self-hosting.

Does it replace V4-Pro? For DeepSeek’s own traffic, effectively yes for now: from September 14, 2026 requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at Flash rates, until a V4.1-Pro ships. On its own tables V4.1 Flash outscores V4-Pro on every row we quote above.

What is its Artificial Analysis Intelligence Index score? There isn’t one yet — AA has not published a score for V4.1 Flash. Our living reports keep the previous build’s point on the chart until it does; we’ll update them when it lands.

Do I have to change anything to keep working? One line: the model id becomes deepseek-v4.1-flash. deepseek-v4-flash still resolves here until October 10, 2026, then returns an invalid-model error. Keys, endpoints, subscriptions and reserved blocks are untouched.

Can it read images? Yes, natively — vision was in the pre-training, not added afterwards. Send image_url parts (OpenAI format) or image blocks (Anthropic format) in user messages, within the 8 MB per-request limit.

Is it better than Claude Opus? On DeepSeek’s own three agent benchmarks above it is ahead of Claude Opus 5.0 — 90.6 vs 89.1, 74.2 vs 74.0, 54.8 vs 50.3. Those are vendor-run numbers at maximum reasoning effort with no independent verification, and two of the three margins are inside a point; treat them as a claim worth testing on your own workload, not a settled ranking.

Full details: DeepSeek V4.1 Flash API docs · Plans & Limits · what a flat monthly DeepSeek subscription changes.

Check live Core Pool availability →


CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.3 and MiniMax M3 (Frontier Pool) and DeepSeek V4.1 Flash and MiMo V2.6 Flash (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

GLM-5.3-Flash: specs, benchmarks, pricing & API options

GLM-5.3-Flash is Z.ai’s (Zhipu AI) efficiency-tier release next to GLM-5.3, and the headline is simple: Artificial Analysis measures it at 57 on its Intelligence Index at a blended price of $0.10 per 1M tokens — $0.09 per Index task, on AA’s intelligence-vs-cost Pareto frontier, ranked #4 of 111 open-weights models it tracks. For scale: the entire index currently tops out at 63.

Unlike its big sibling — a post-training pass on the existing 743B GLM base — Flash is a new model: 320B total / 18B active MoE with hybrid sparse + linear attention, trained on a 30T-token multimodal corpus, and natively multimodal (image, video and file input) where GLM-5.3 is text-only. The weights are on Hugging Face under MIT (zai-org/GLM-5.3-Flash) — plain MIT, where the flagship’s custom license carries a security-review clause for large Model-as-a-Service operators.

Architecture320B total / 18B active MoE, hybrid sparse + linear attention
MultimodalNative — image, video, text and file input; text output
Context window1M tokens; up to 128K output
ReasoningReasoning model — thinking always on, cannot be disabled on the direct API
Open weightsYes — zai-org/GLM-5.3-Flash, MIT license (BF16 + FP8)
List price (API)$0.15 in / $0.50 out per 1M, cached input $0.03 — 50% launch promo ($0.075 / $0.25) until September 9, 2026
Measured speed48.6 output tok/s, 1.51s to first token (Artificial Analysis)

Z.ai’s own reported numbers (vendor harness — not independently reproduced), with the retired GLM 5.2 and the full GLM-5.3 on either side:

GLM 5.2 (743B, retired)GLM-5.3-Flash (18B active)GLM-5.3 (743B)
Terminal-Bench 2.1 DeepSWE v1.1 AutomationBench GLM 5.2 — Terminal-Bench 2.1: 81.0 GLM-5.3-Flash — Terminal-Bench 2.1: 84.3 GLM-5.3 — Terminal-Bench 2.1: 88.2 GLM 5.2 — DeepSWE v1.1: 46.2 GLM-5.3-Flash — DeepSWE v1.1: 63.4 GLM-5.3 — DeepSWE v1.1: 66.9 GLM 5.2 — AutomationBench v1.0.6: 26.2 GLM-5.3-Flash — AutomationBench v1.0.6: 48.8 GLM-5.3 — AutomationBench v1.0.6: 48.2 81.084.388.2 46.263.466.9 26.248.848.2
Benchmark (Z.ai, vendor-run)GLM 5.2GLM-5.3-FlashGLM-5.3
Terminal-Bench 2.181.084.388.2
DeepSWE v1.146.263.466.9
AutomationBench v1.0.626.248.848.2

Two things stand out. An 18B-active model beats the 743B GLM 5.2 on all three — a model that was frontier-class until weeks ago. And on AutomationBench it edges out the full GLM-5.3 itself. Z.ai also reports Flash within half a point of Claude Opus 4.8 on its internal Code Bench v1.0 (29.0 vs 29.5, max effort) — vendor-run, so hold it loosely until independent reproductions land.

The independent signal, from Artificial Analysis (current index edition — every score below re-verified today):

GLM-5.3-FlashServed in a CheapestInference poolReference frontier models
Claude Opus 5 Claude Fable 5 GPT-5.6 Sol Kimi K3 GLM-5.3 Qwen3.8 Max GLM-5.3-Flash Claude Opus 5 (max effort): 63 — current index leader Claude Fable 5 (max effort, Opus 4.8 fallback — AA's evaluated config): 62 GPT-5.6 Sol (max): 61 Kimi K3 (max) — Flagship Pool: 60, tied top open-weights score GLM-5.3 (max) — Frontier Pool: 60, tied top open-weights score — blended $0.90/1M (AA) Qwen3.8 Max — Flagship Pool: 58 GLM-5.3-Flash: 57 — blended $0.10/1M (AA), #4 open-weights model on the index 6362 6160 6058 57

Honest framing, as always: Flash does not beat the closed frontier — Claude Opus 5 leads the index at 63 — and it is three points behind the best open-weights scores (Kimi K3 and GLM-5.3, tied at 60). What is remarkable is the column AA puts next to those scores:

ModelAA Intelligence IndexAA blended $/1M tokens
GLM-5.3 (max)60$0.90
Qwen3.8 Max58—
GLM-5.3-Flash57$0.10

Blended = AA’s 7:2:1 cache-hit/input/output mix; scores from the current index edition. Full field and history in our monthly LLM Pareto Frontier report.

95% of GLM-5.3’s measured intelligence at a ninth of its blended price is why AA places Flash on the Pareto frontier: among everything it tracks at this intelligence level, nothing is cheaper per task. The trade-off it does not hide: speed. Flash generates ~48.6 tok/s against GLM-5.3’s ~66.6, so interactive latency is where the smaller model feels smaller.

Z.ai shipped GLM-5.3’s weights under a custom license with a security-review clause for Model-as-a-Service operators above US$10B revenue. Flash ships under plain MIT — no clauses, BF16 and FP8 checkpoints on Hugging Face. For anyone evaluating models to serve rather than just call, that difference is not cosmetic: MIT is as serving-friendly as licenses get, and it makes Flash the most permissively-licensed near-frontier model of the moment.

We serve the full GLM-5.3 unlimited in the Frontier Pool (from $71/mo, flat) since August 30. A near-frontier, MIT-licensed, 1M-context multimodal model is squarely the profile our pools exist for, and Flash is under active evaluation — the live pipeline status is always on the models-under-review page, and additions land in the changelog the day they ship.

What is GLM-5.3-Flash? Z.ai’s efficiency-tier model beside GLM-5.3: a new 320B-total / 18B-active MoE with native multimodal input (image, video, file), a 1M-token context window and MIT-licensed open weights. Artificial Analysis scores it 57 on its Intelligence Index at $0.10 per 1M tokens blended.

How much does the GLM-5.3-Flash API cost? Z.ai lists $0.15 input / $0.50 output per 1M tokens (cached input $0.03), with a 50% launch discount until September 9, 2026. Artificial Analysis measures $0.09 per Intelligence-Index task and a $0.10/1M blended rate.

Does GLM-5.3-Flash have open weights? Yes — zai-org/GLM-5.3-Flash on Hugging Face under the MIT license, in BF16 and FP8. More permissive than the full GLM-5.3, whose custom license adds a security-review clause for large Model-as-a-Service operators.

How good is GLM-5.3-Flash at coding? On Z.ai’s vendor benchmarks it scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1 — above the retired 743B GLM 5.2 on both — and 48.8 on AutomationBench, marginally above the full GLM-5.3. Independently, its 57 on the AA index is three points behind the best open-weights models. Its measured output speed is ~48.6 tok/s.

Is there an unlimited GLM-5.3-Flash API subscription? Not from us today — Flash is under evaluation (live status). The full GLM-5.3 is served unlimited in the Frontier Pool from $71/mo: flat monthly fee, no token caps during your reserved hours.


CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.3 and MiniMax M3 (Frontier Pool) and DeepSeek V4.1 Flash and MiMo V2.6 Flash (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Claude Code alternatives in 2026: switch tools, or switch the model behind it

Searches for a “Claude Code alternative” usually start with one of two pains: the subscription limits (the five-hour usage windows belong to the Claude plan, not to the tool) or the bill (agentic coding burns tokens like nothing else you run). Before comparing tools, it helps to split the question in two, because they have different answers:

  • The harness — the terminal agent itself: the REPL, the tool-calling loop, permissions, MCP support. Claude Code is one harness; there are now several good ones.
  • The model — what actually writes the code, and what you’re actually paying for.

You can swap either one independently. Here’s the honest map of both.

The real alternatives (swapping the harness)

Section titled “The real alternatives (swapping the harness)”
ToolByModels it drivesWorth knowing
Codex CLIOpenAIGPT familyShips with ChatGPT plans; open-source harness
Gemini CLIGoogleGemini familyGenerous free tier; the harness many forks build on
OpenCodeSSTAny (bring your own API)Open-source, provider-agnostic, closest to Claude Code in feel
Qwen CodeAlibabaQwen family + any OpenAI-compatible APIGemini CLI fork tuned for Qwen
Aideropen-sourceAny (bring your own API)The veteran; git-native, great diffs
Cline / Roo Codeopen-sourceAny (bring your own API)VS Code sidebar instead of a terminal

All of these are good software. If your pain is the harness itself — you want an IDE sidebar, or a different permission model — pick from the table and you’re done. But notice what the table also says: half of these tools don’t come with a model at all. You bring an API key, and the model behind it is where the cost and the quality actually live.

The option most people miss: keep Claude Code, swap the model

Section titled “The option most people miss: keep Claude Code, swap the model”

Claude Code talks to any endpoint that implements the Anthropic Messages API — that’s two environment variables:

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-..."
export ANTHROPIC_MODEL="kimi-k3" # or glm-5.3, deepseek-v4.1-flash, ...

Your muscle memory, your .claude/ config, your MCP servers, your slash commands — everything stays. What changes is the engine and the meter: on a CheapestInference pool, usage during your reserved time blocks is a flat monthly fee, with no per-token billing and no five-hour windows.

Model-by-model setup guides:

How the per-token prices of these models compare to the closed frontier is tracked monthly in the LLM API pricing comparison.

If your only pain is “usage limit reached”

Section titled “If your only pain is “usage limit reached””

If you’re otherwise happy on a Claude subscription and just want overnight runs to survive the window resets, you don’t need to switch anything: claude-auto-retry waits out the limit and resumes the session for you. Free, one npm install.

The harnesses are mostly free and open-source (OpenCode, Aider, Cline, Qwen Code, Codex CLI’s source). What’s never free at scale is the model behind them — a serious agentic session runs millions of tokens, so the real comparison is per-token bills vs. subscriptions vs. flat-rate blocks.

Can Claude Code use models other than Claude?

Section titled “Can Claude Code use models other than Claude?”

Yes. Claude Code works with any Anthropic-compatible endpoint via ANTHROPIC_BASE_URL — no plugin, no proxy. That’s the “keep the harness, swap the model” path above.

What’s the cheapest way to run a coding agent all day?

Section titled “What’s the cheapest way to run a coding agent all day?”

A flat-rate block: agentic coding is exactly the workload where per-token billing hurts most, because the agent — not you — decides how many tokens to spend. A pool subscription makes the heavy week cost the same as the light one. For per-token numbers across the market, see the live pricing comparison.