Skip to content

guides

7 posts with the tag “guides”

DeepSeek V4-Flash-0731: what changed, and how to use it over the API

On July 31, DeepSeek shipped V4-Flash-0731 — not a bigger model, a retrained one. Same 284B-total / 13B-active architecture, same 1M-token context, same price bracket — but re-post-trained for agent work, and the reported jump is unusual: DeepSeek says the new Flash now beats its own larger V4-Pro-Preview on every one of the nine agent benchmarks it publishes.

If you use the Core Pool, there is nothing to migrate: the model id is still deepseek-v4-flash, and requests already serve the new build. Unlimited, flat-rate access starts from $14.99/mo ($12.74/mo billed annually) — live pricing on /pools.

Same architecture, new post-training. This is DeepSeek’s own reported before/after (vendor harness — no independent verification of these specific numbers yet):

V4-Flash-0731V4-Flash preview (the previous build, April 2026)
Terminal-Bench 2.1 DSBench-FullStack DeepSWE V4-Flash-0731 — Terminal-Bench 2.1: 82.7 V4-Flash preview — Terminal-Bench 2.1: 61.8 V4-Flash-0731 — DSBench-FullStack: 68.7 V4-Flash preview — DSBench-FullStack: 37.0 V4-Flash-0731 — DeepSWE: 54.4 V4-Flash preview — DeepSWE: 7.3 82.761.8 68.737.0 54.47.3

Full vendor-stated table for the 0731 build, next to the April preview it replaces:

BenchmarkFlash-0731Flash preview (Apr)
Terminal-Bench 2.182.761.8
Cybergym76.738.7
Toolathlon-Verified70.349.7
DSBench-FullStack†68.737.0
DSBench-Hard†59.625.8
DeepSWE54.47.3
NL2Repo54.239.4
Agents’ Last Exam25.215.8
AutomationBench Public25.110.8

† DeepSeek-internal test sets; the rest are public benchmarks.

The DeepSWE number is the striking one: the retrain multiplied the preview’s score by seven without touching the architecture. And like the April preview, the 0731 checkpoint is a genuine open-weights release: the weights are published under MIT at deepseek-ai/DeepSeek-V4-Flash-0731 (304B parameters on the repo — the 284B base plus a draft module).

The independent signal comes from Artificial Analysis, which measures Flash-0731 at 50 on its Intelligence Index — ten points above the previous Flash. Honest framing: it does not beat Claude Opus 5 (61) or GPT-5.6 Sol (59) — nothing near this price does. What it does is land within seven points of Kimi K3 (57, the highest open-weights score on the index) from the budget model of the field — with a reference per-token price of $0.28 per 1M output tokens, roughly 90× below Claude Opus 5’s $25:

Served in a CheapestInference poolReference frontier models
Claude Opus 5 Claude Fable 5 GPT-5.6 Sol Kimi K3 Claude Opus 4.8 GLM 5.2 V4-Flash-0731 MiniMax M3 Claude Opus 5 (max effort): 61 — list output $25.00/1M Claude Fable 5 (max effort, Opus 4.8-fallback config as evaluated by AA): 60 — list output $50.00/1M GPT-5.6 Sol (max): 59 — list output $30.00/1M Kimi K3 (max) — Flagship Pool: 57 — list output $15.00/1M Claude Opus 4.8: 56 — list output $25.00/1M GLM 5.2 (max) — Frontier Pool: 51 — list output $4.40/1M DeepSeek V4-Flash-0731 — Core Pool: 50 — list output $0.28/1M MiniMax M3 — Frontier Pool: 44 — list output $1.20/1M 6160 5957 5651 5044

Same data as a table, with list output prices alongside for scale — each model at its best published configuration (max effort where the index reports one); the median across all models Artificial Analysis tracks is 17:

ModelAA Intelligence IndexList output $/1MOn CheapestInference
Claude Opus 561$25.00
Claude Fable 5*60$50.00
GPT-5.6 Sol59$30.00
Kimi K357$15.00Flagship Pool
Claude Opus 4.856$25.00
GLM 5.251$4.40Frontier Pool
DeepSeek V4-Flash-073150$0.28Core Pool
MiniMax M344$1.20Frontier Pool

* Yes, one point below Claude Opus 5, even though Anthropic positions Fable 5 above Opus in capability. The index scores what AA actually evaluates: Fable 5 is measured in its “Adaptive Reasoning, Max Effort, Opus 4.8 Fallback” serving configuration — the variant with Anthropic’s additional dual-use safeguards — while Opus 5 runs at plain max effort. Independent index, published configurations, taken as-is.

Four of the eight models on that chart are served here on flat-rate subscriptions — and the 0731 retrain just moved the cheapest pool of the three into frontier territory.

V4-Flash-0731 vs Claude Opus 4.8, GPT-5.6 Luna, and Gemini 3.6 Flash

Section titled “V4-Flash-0731 vs Claude Opus 4.8, GPT-5.6 Luna, and Gemini 3.6 Flash”

Is DeepSeek V4-Flash-0731 better than Claude Opus? On raw capability, no — and we won’t pretend otherwise. The strongest competitor in DeepSeek’s own release table is Claude Opus 4.8, and Opus wins every one of the nine shared benchmarks. What the table actually shows is how little it wins by, against a model priced roughly 90× higher per output token at list:

Benchmark (DeepSeek’s release table, vendor-run)V4-Flash-0731Claude Opus 4.8
Terminal-Bench 2.182.785.0
Cybergym76.783.1
Toolathlon-Verified70.376.2
DSBench-FullStack†68.771.6
DSBench-Hard†59.671.7
DeepSWE54.458.0
NL2Repo54.269.7
Agents’ Last Exam25.225.7
AutomationBench Public25.127.2
Reference output price, per 1M tokens$0.28$25.00

On Agents’ Last Exam the gap is half a point — effective parity. On Terminal-Bench 2.1 it is 2.3 points. NL2Repo and the (internal) DSBench-Hard are where the distance stays wide.

Against its actual price peers, the independent picture flips. Per Artificial Analysis: Flash-0731 sits one point behind GPT-5.6 Luna (50 vs 51 at max effort) with a cost per task roughly 60% lower — even after OpenAI’s price cut — ties Gemini 3.6 Flash (50), and lands on AA’s Pareto frontier for Intelligence vs Cost per Task: at this intelligence level, nothing tracked is cheaper per task.

Here is that frontier drawn out — intelligence against list output price. A model is on the frontier when nothing tracked is both smarter and cheaper; everything below-right of the line pays more for less. The 0731 retrain moved Flash onto it, and pushed MiniMax M3 off:

Served in a CheapestInference poolReference frontier modelsPareto frontier
4045 5055 60 $0$10 $20$30 $40$50 List output price — $ per 1M tokens AA Intelligence Index ↑ DeepSeek V4-Flash-0731 — Core Pool: index 50 at $0.28/1M — on the frontier MiniMax M3 — Frontier Pool: index 44 at $1.20/1M GLM 5.2 — Frontier Pool: index 51 at $4.40/1M — on the frontier Gemini 3.5 Flash: index 50 at $9.00/1M Claude Sonnet 5: index 53 at $10.00/1M — on the frontier Kimi K3 — Flagship Pool: index 57 at $15.00/1M — on the frontier Claude Opus 4.8: index 56 at $25.00/1M Claude Opus 5: index 61 at $25.00/1M — on the frontier GPT-5.6 Sol: index 59 at $30.00/1M Claude Fable 5: index 60 at $50.00/1M (AA config: max effort, Opus 4.8 fallback) V4-Flash-0731 MiniMax M3 GLM 5.2 Gemini 3.5 Flash Sonnet 5 Kimi K3 Opus 4.8 Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Flash-0731 matches Gemini-Flash-class intelligence at 32× lower list price

Three of the five models on that frontier — Flash-0731, GLM 5.2 and Kimi K3 — are served here on flat rate. This chart is a snapshot of the frontier’s cheap end; the monthly-updated, full-field version (with cost-per-task data and edition history) lives in our LLM Pareto Frontier report.

And on a flat-rate subscription the per-token column stops mattering altogether: a Core Pool block is the same price whether your agent burns one million tokens or one billion.

curl:

Terminal window
curl https://api.cheapestinference.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek-v4-flash", "messages": [{"role": "user", "content": "Hello"}]}'

OpenAI SDK (Python):

from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="sk-...", # your subscriber key
)
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Fix the failing test in this repo..."}],
)
print(response.choices[0].message.content)

No account yet? Register, subscribe to the Core Pool, and mint a key at cheapestinference.com/keys.

The API also speaks the Anthropic Messages format, so a retrained agent model drops straight into the most popular coding agent:

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-..." # your subscriber key
export ANTHROPIC_MODEL="deepseek-v4-flash"
export ANTHROPIC_SMALL_FAST_MODEL="deepseek-v4-flash"

Start claude as usual — every request runs on Flash-0731 with no per-token meter. The same key works in Cline, Roo Code, Continue, and anything that accepts a custom OpenAI base URL (setup guides).

Why this release matters for flat-rate users

Section titled “Why this release matters for flat-rate users”

Agent workloads are exactly where per-token bills explode: an agent re-sends its growing context on every tool call, and iteration count — not task value — drives the invoice. A model that is suddenly much better at agent work makes that math worse per-token and better flat-rate: more capable loops, same fixed monthly price. We ran the full per-token vs. flat break-even math in Unlimited DeepSeek: what a flat monthly subscription changes — every row of that table just got more favorable, because the same subscription now serves a stronger model.

The trade-offs, as always: usage is unlimited in tokens during your reserved 8-hour blocks, each subscription runs one request at a time (fair use), and outside your blocks the key doesn’t serve. Full details: DeepSeek V4 Flash API docs · Plans & Limits.

Check live Core Pool availability →


CheapestInference serves Kimi K3 (Flagship Pool), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

How to use the Kimi K3 API: key, snippets, and unlimited access

Kimi K3 is live on CheapestInference: Moonshot’s flagship — scoring 57 on the Artificial Analysis Intelligence Index, just 4 points off the top model and the narrowest open-vs-closed gap since February — served through an OpenAI- and Anthropic-compatible API with unlimited usage at a flat monthly price, from $149/mo (live availability — seats are very limited). This is the practical guide: get access, call it, wire it into your coding agent.

Three realistic routes to K3 over an API today:

  1. Per-token, from Moonshot — $3.00 per 1M input tokens ($0.30 cached) and $15.00 per 1M output. Ideal for evaluation and light use; expensive fast for agent workloads.
  2. Kimi memberships — Moonshot’s own plans bundle K3 access with request quotas per 5-hour and weekly windows (new signups have been intermittently paused since launch).
  3. Unlimited time-block subscription (this guide) — reserve one or more daily 8-hour blocks on the Flagship Pool and use K3 with no token caps during your hours. From $149/mo per block; all three blocks = 24/7.

For the subscription route: create an account, subscribe to the Flagship Pool, and mint an API key at cheapestinference.com/keys. Model id: kimi-k3.

curl:

Terminal window
curl https://api.cheapestinference.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "kimi-k3", "messages": [{"role": "user", "content": "Hello"}]}'

OpenAI SDK (Python):

from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="sk-...", # your subscriber key
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Refactor this function..."}],
)
print(response.choices[0].message.content)

K3 is a reasoning model, but reasoning is off by default — you get fast, direct answers. Turn it on per request with "thinking": {"type": "enabled"} or reasoning_effort (low / high / max — K3’s deepest tier). With reasoning on, give it a generous max_tokens so long answers don’t truncate mid-thought, and stream (stream: true) for responsive UIs.

The API also speaks the Anthropic Messages format, so Claude Code works out of the box:

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-..." # your subscriber key
export ANTHROPIC_MODEL="kimi-k3"
export ANTHROPIC_SMALL_FAST_MODEL="kimi-k3"

Start claude as usual — every request now runs on K3 with no per-token meter. The same key works in Cline, Roo Code, Continue, and any client that accepts a custom OpenAI base URL (setup guides).

At list price, a coding agent that burns 100M tokens a month on K3 costs roughly $200–400 per-token, depending on your cache-hit and output mix — and heavy agentic users go far beyond that. A Flagship block at $149/mo covers your working day; all three blocks cover 24/7. The trade-offs: during your reserved hours usage is truly unlimited, each subscription runs one request at a time (fair use), and outside your blocks the key doesn’t serve. Full pricing detail: Plans & Limits · Kimi K3 API docs.

Seats on the Flagship Pool are very limited — when a block sells out it’s gone until someone leaves. Check live availability →


CheapestInference serves Kimi K3 (Flagship Pool), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Combined keys: stack your subscriptions into one credential

On an Unlimited subscription, the thing you’re actually shaping isn’t tokens — it’s capacity over time. Each subscription gives you a daily coverage window (the 8-hour blocks you reserve) and, within it, some number of requests you can run in parallel. Your real capacity is those two dimensions multiplied: parallel slots × hours.

Most people grow along both axes. You buy a second block to cover more of the day, or a second subscription to run more requests at once during your busy hours. The awkward part used to be the bookkeeping: every subscription minted its own key, so a serious setup meant three or four keys, each live at different times, each with its own capacity — and you juggling which one to paste where.

Combined keys remove the juggling. A combined key (it looks like sk-ci-meta-…) folds two or more of your subscriptions into a single credential, and it behaves like the union of everything underneath it.

Coverage adds up. A combined key is live whenever any of its subscriptions has an open window. Reserve the Europe block on one subscription and the Americas block on another, combine them, and the one key covers both stretches of the day.

Overlap stacks parallel capacity. Where two subscriptions cover the same hour, their parallel slots add together. That’s the lever for concurrency: if you need to run more requests side by side during your peak hours, buy a second subscription over those hours and combine it in.

A concrete example. Say you hold one full-day (24h) subscription and add a second subscription on just the Europe block. The combined key gives you:

  • Double the parallel capacity during the Europe hours — the two subscriptions overlap there, so their slots stack.
  • Baseline capacity the rest of the day — only the 24h subscription is covering those hours.

Coverage is 24/7 (from the full-day subscription); parallel capacity is shaped to peak exactly when you work.

Combining is about capacity, not billing. Each subscription keeps its own monthly allowance and its own renewal date — nothing is pooled or co-mingled. The practical upside is resilience: if one subscription lapses or you let it cancel, it simply drops out of the union. The combined key keeps working with whatever subscriptions remain — no dead key, no scramble to re-issue credentials.

Requests route by the model you ask for. So a combined key can span subscriptions on different pools — a Core Pool subscription and a Frontier Pool subscription under one key — and each request lands wherever its model lives. Ask for deepseek-v4-flash and it’s served from your Core subscription; ask for kimi-k2.7 and it’s served from your Frontier one. GET /v1/models on a combined key lists every model you can reach across all of them.

Combining a subscription into a key removes that subscription’s standalone key — each subscription has exactly one credential at a time. That’s deliberate: it means there’s never ambiguity about what capacity a given key carries. A key’s coverage and parallel capacity are always the exact sum of the subscriptions it holds, nothing more, nothing hidden. The dashboard walks you through it when you combine.

In the Keys page, choose Create API Key and multi-select the subscriptions you want to combine. Before you commit, a preview shows the resulting daily coverage and peak parallel capacity, so you can see the shape you’re buying into. Prefer the API? The Management API does the same thing programmatically.

Full walkthrough in the combined keys guide. If you’re still deciding which blocks to reserve, the live menu and per-block prices are on the pools page.


CheapestInference serves frontier open-weights models — Kimi K2.7, GLM 5.2, MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash, MiMo v2.5 (Core Pool) — through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Claude usage limit reached? How to auto-resume Claude Code

It’s 11pm. Claude Code is 40 minutes into a refactor, tests are half-green, and then the terminal prints:

5-hour limit reached ∙ resets 3am

The agent stops. The session sits there. If you go to bed, nothing happens at 3am — the limit resets, but nobody is around to type “continue”. You lose the hours between the reset and whenever you come back.

We got tired of this, so we built claude-auto-retry: a small open-source tool that watches your Claude Code session, parses the reset time out of the limit message, waits for it, and resumes the session automatically. You run claude exactly like before. That’s the whole pitch.


The problem: limits don’t care where your task is

Section titled “The problem: limits don’t care where your task is”

Claude subscriptions meter usage in rolling windows. Heavy Claude Code use — long agent runs, big contexts, lots of tool calls — burns through a window fast, and when it’s gone you get a variant of:

  • 5-hour limit reached ∙ resets 3pm
  • You've hit your limit · resets 3pm
  • Claude usage limit reached. Resets at 2pm

The limit itself is fine — it’s how a fixed-price subscription stays fixed-price. The annoying part is what happens after: Claude Code doesn’t queue your work and resume at the reset. It just stops, mid-task, holding all the context of whatever it was doing. Resuming costs one word — “continue” — but only if a human is there to type it.

For interactive use that’s a coffee break. For the way people increasingly use Claude Code — overnight runs, long autonomous tasks, claude -p in scripts — it’s a silent halt that wastes every hour between the reset and your return.

Workarounds people try (and why they’re fragile)

Section titled “Workarounds people try (and why they’re fragile)”

Sitting there. Works, wastes an evening.

A cron job that types “continue” at a fixed time. Blind. The reset time moves with your usage, so you either fire too early (message goes nowhere, session still limited) or too late (hours lost anyway). And if Claude exited, you’re injecting keystrokes into a bare shell.

Wrapper scripts around the CLI. Most break something: interactive mode, claude -p piping, or they lose the session when your terminal disconnects.

The failure modes all reduce to the same two hard parts: knowing when the limit actually resets, and injecting “continue” safely into a live session — including after you’ve closed your laptop.

The tool leans on tmux for session persistence, but hides it completely:

  1. Transparent wrapping. The installer adds a shell function so claude transparently runs inside a tmux session. If you’re already in tmux, it uses your pane. You won’t notice the difference — interactive use, flags, and claude -p "prompt" | jq all work as before.

  2. Background monitoring. A separate process polls the pane every 5 seconds looking for the known limit-message patterns (customizable if Anthropic changes the wording).

  3. Real reset-time parsing. When a limit hits, it parses the actual reset time from the message — timezone-aware and DST-safe — computes the wait, and sleeps until reset plus a small safety margin. No guessed schedules.

  4. Safe injection. Before sending “continue”, it verifies Claude is still the foreground process in the pane. If Claude exited, nothing gets typed into your shell.

  5. Survives disconnects. Because the session lives in tmux, you can close the terminal, drop SSH, or shut the laptop lid. The monitor keeps waiting server-side; reattach whenever and the run has continued without you.

Separately from usage limits, it also catches transient API errors (429, 5xx, 529 overloaded) and retries those with exponential backoff — a different failure with a different fix, handled by the same monitor.

It’s pure Node.js with zero npm dependencies, MIT-licensed, and logs everything it does to ~/.claude-auto-retry/logs/.

Terminal window
npm i -g claude-auto-retry
claude-auto-retry install

The installer adds the shell function to ~/.bashrc or ~/.zshrc and installs tmux if it’s missing (apt, dnf, brew, pacman, apk). Then use Claude Code exactly as always:

Terminal window
claude

When a limit hits, you’ll see the monitor take over. Check on it with:

Terminal window
claude-auto-retry status # what it's waiting for, and until when
claude-auto-retry logs # what it has done

Retry counts, the wait margin, and custom detection patterns live in ~/.claude-auto-retry.json if you want to tune them. Needs Node.js ≥ 18; tested on Ubuntu, CentOS, macOS, Arch, and Alpine.

We first shared the tool on r/ClaudeAI — the discussion there is a decent picture of how many people hit this exact wall.

How do I resume Claude Code after hitting the usage limit? Wait for the reset time shown in the limit message, then type continue in the same session — Claude Code keeps all its context. To do it automatically, claude-auto-retry parses the reset time from the message, waits, and types continue for you.

Can Claude Code auto-continue after the limit resets? Not on its own — Claude Code stops mid-task and waits for a human. That’s exactly the gap claude-auto-retry fills: one npm i -g claude-auto-retry, and sessions resume at the reset without anyone at the keyboard.

Does Claude Code have a 5-hour limit? The limit belongs to your Claude subscription, not to Claude Code itself: usage is metered in rolling windows shared across everything on the account, and heavy agent runs simply burn the window faster. When it’s exhausted, every surface — Claude Code included — waits for the reset.

Can Claude Code resume overnight, with my laptop closed? Yes, if the session survives your terminal: claude-auto-retry runs Claude Code inside tmux transparently, so you can drop SSH or close the lid and the monitor still types continue at the reset, server-side. Reattach in the morning to a run that kept going.

Can I avoid the usage limit entirely for unattended runs? Not on a metered subscription — but you can route the unattended work to an endpoint without rolling caps. See the next section.


claude-auto-retry makes limits painless when the work can wait a few hours. Some workloads can’t — always-on agents, batch jobs, anything that has to keep moving at 4am.

Claude Code speaks the Anthropic Messages API, which means it can point at any compatible endpoint — including ours. CheapestInference serves open-weight models across three pools (Kimi K3 in the Flagship pool; Kimi K2.7, GLM 5.2, and MiniMax M3 in the Frontier pool; DeepSeek V4 Flash and MiMo v2.5 in the budget Core pool) through an Anthropic-compatible endpoint, with a different limits model: you reserve daily 8-hour time blocks, and during your reserved hours inference is unlimited — there is no usage-limit message to parse, because there is no rolling cap to hit.

Terminal window
export ANTHROPIC_BASE_URL="https://api.cheapestinference.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="your-api-key"
export ANTHROPIC_MODEL="kimi-k2.7"
claude

The two approaches compose. Plenty of people keep their Claude subscription for interactive work and run the long unattended stuff — the runs that would otherwise die at 11pm — on an open-weight model during a reserved block. Open-weight coding models are closer to frontier than most people assume; here’s the current evidence.

Either way: stop typing “continue”.

Self-hosted vs. API inference: the real cost comparison

“Why pay for an API when I can run the model myself?”

It’s a reasonable question. Open-source models are free. GPUs are available on every cloud. vLLM and Ollama make serving straightforward. The math should be simple: GPU cost per hour × hours = total cost. Done.

Except it’s not. The GPU is the minority of the cost. Here’s the full picture.


Running DeepSeek V3.2 (671B MoE, ~130B active parameters) requires at least 4× A100 80GB or 2× H100 80GB in FP8. Qwen 3.5 397B has similar requirements.

Setup Hourly Monthly (24/7) Monthly (8h/day)
4× A100 80GB (cloud) $12.80 $9,216 $2,816
2× H100 80GB (cloud) $8.40 $6,048 $1,848
1× A100 80GB (Llama 70B) $3.20 $2,304 $704
1× L40S (Llama 8B) $1.10 $792 $242

These are cloud GPU rental prices (AWS, GCP, Lambda Labs — varies by provider and availability). If you buy hardware, the upfront cost is $15K–$40K per GPU, amortized over 3–4 years, plus electricity, cooling, and data center costs.

Smaller models are cheaper — but limited

Section titled “Smaller models are cheaper — but limited”

Running Llama 3.1 8B on a single L40S costs $242/month (8h/day). That’s competitive with API pricing. But 8B models can’t handle complex coding, multi-step reasoning, or nuanced analysis — the tasks where AI provides the most value.

The models worth self-hosting (70B+, MoE) require multi-GPU setups where the economics change dramatically.


GPU rental is just the beginning.

Someone has to:

  • Set up vLLM/TGI with optimal batch sizes, quantization, and memory allocation
  • Monitor GPU utilization and restart crashed processes
  • Update model weights when new versions release
  • Handle OOM errors, NCCL failures, and driver issues
  • Manage the serving infrastructure (load balancer, health checks, auto-scaling)

If this is a full-time DevOps engineer at $150K/year, that’s $12,500/month in labor. If it’s 20% of a senior engineer’s time, it’s $2,500/month. Either way, it’s more than the GPU.

GPUs cost money whether they’re inferring or not. If your usage pattern is 8 hours of heavy use (work hours) and 16 hours of near-zero traffic, you’re paying for 24 hours and using 8.

Cloud spot instances help but introduce availability risk. Auto-scaling GPU clusters is possible but complex — model loading takes minutes, not seconds.

API pricing is purely usage-based. Zero requests = zero cost.

Self-hosting one model is manageable. Self-hosting five models for different tasks — a coding model, a reasoning model, a fast classification model, an embedding model, and a vision model — requires either:

  • 5 separate GPU instances (expensive)
  • Shared GPU with model swapping (slow — loading a 70B model takes 2–5 minutes)
  • A serving framework that handles multi-model routing (complex)

An API gives you access to many models through the same endpoint. No model loading, no GPU allocation, no routing logic.

Every hour your team spends on inference infrastructure is an hour not spent on your actual product. For startups, this is the most expensive cost of all — it doesn’t show up on any invoice.


For a team of 5 developers running AI-assisted coding with a mix of DeepSeek V3.2 and smaller models:

Cost Self-hosted API (per-token) API (time-block sub)
Compute/inference $2,800 $265 $250
Ops/maintenance $2,500 $0 $0
Idle waste (~60%) $1,680 $0 $0
Total monthly $6,980 $265 $250

Self-hosting costs 26x more for the same workload. The GPU is only 40% of the self-hosted cost — ops and idle waste are the majority.


Self-hosting wins in specific scenarios:

Data sovereignty: If your data cannot leave your network — regulated industries, government, healthcare with strict compliance — self-hosting is the only option. No API provider can guarantee the data isolation you need.

Extreme scale: If you’re processing millions of requests per day and your GPUs are consistently at 80%+ utilization, the per-token math eventually favors owned hardware. This threshold is higher than most teams expect — typically $20K+/month in API spend before self-hosting breaks even.

Custom models: If you’ve fine-tuned a model and need to serve it, self-hosting or a dedicated inference provider (Fireworks, Together) is required. Most unified APIs don’t serve custom model weights.

Latency control: If you need guaranteed sub-100ms TTFT and your data center is co-located with your GPUs, self-hosting eliminates network hops.

For everyone else — startups, small teams, companies with variable usage patterns — the API is cheaper, faster to set up, and easier to maintain.


Most teams don’t need to choose one forever. A practical approach:

  1. Start with an API: Get your product working, validate demand, understand your usage patterns.
  2. Optimize model selection: Use cheaper models for simple tasks, frontier models for hard tasks. Full guide: Multi-model architecture.
  3. Evaluate self-hosting when: Your monthly API spend exceeds $10K, your GPU utilization would be >70%, and you have DevOps capacity to maintain it.
  4. Hybrid: Self-host your high-volume models, use an API for long-tail models and overflow capacity.

The worst outcome is spending 3 months setting up GPU infrastructure before you’ve validated that anyone wants your product.


CheapestInference serves five open-weight models across two pools — Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier); DeepSeek V4 Flash and MiMo v2.5 (Core) — through a single API. No GPUs to manage, no idle costs, no ops burden. Reserve a daily 8-hour time block for unlimited usage from $12.74/mo (Core) or $48.45/mo (Frontier), billed annually — reserve all three blocks for full 24/7. Get started or see the pools.

Building a multi-model architecture: route requests to the right LLM

Using one model for everything is the simplest architecture. It’s also the most wasteful. A 685B-parameter reasoning model answering “what’s the weather?” is like hiring a PhD to sort mail.

This guide covers how to use a small, fast model to classify incoming requests and route them to the right specialist. The result: lower latency, lower cost, and often better quality — because each model handles what it’s actually good at.


The problem with single-model architectures

Section titled “The problem with single-model architectures”

Most applications start with one model:

User request --> Large Model --> Response

This works, but every request — simple or complex — pays the same latency and cost penalty. When 60% of your traffic is simple classification, FAQ, or extraction, you’re burning expensive compute on tasks a small model handles equally well.

Llama 3.1 8B
~200 t/s
DeepSeek V3.2
~60 t/s
DeepSeek R1
~30 t/s

The gap between Llama 8B and R1 is nearly 7x in throughput. Routing simple requests to the small model saves that difference on every request.


User request --> Router (GLM 5.2) --> classify intent
|
+-----------+-----------+-----------+
| | | |
simple/general reasoning code agent
| | | |
GLM 5.2 MiniMax M3 MiniMax M3 Kimi K2.7
| | | |
+-----+-----+-----+-----+
|
Response

Two stages:

  1. Classify — The router model reads the user’s message and outputs a category. A fast model returns this in a fraction of a second.
  2. Route — Based on the category, forward the request to the appropriate specialist model.

The router adds minimal overhead (~200ms) but saves significant compute by keeping simple requests away from expensive models.


A fast, lightweight model makes a good router. With low TTFT and a short, single-word output, the classification step costs almost nothing and completes before the user notices. (On CheapestInference, DeepSeek V4 Flash or MiMo v2.5 in the Core pool are natural router models; the example below uses GLM 5.2 so everything runs on one Frontier subscription.)

The classification prompt is simple — you want a single-word category, not a conversation:

from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="your-api-key"
)
def classify_request(user_message: str) -> str:
"""Classify a user message into a routing category."""
response = client.chat.completions.create(
model="glm-5.2",
messages=[
{
"role": "system",
"content": (
"Classify the user's message into exactly one category. "
"Respond with only the category name, nothing else.\n\n"
"Categories:\n"
"- simple: greetings, FAQ, simple factual questions\n"
"- general: complex questions, analysis, writing, summarization\n"
"- reasoning: math, logic, multi-step problems, science\n"
"- code: code generation, debugging, refactoring, technical implementation\n"
"- agent: tasks requiring tool use, web search, or multi-step execution"
)
},
{"role": "user", "content": user_message}
],
max_tokens=10,
temperature=0
)
category = response.choices[0].message.content.strip().lower()
# Default to general if classification is unclear
valid = {"simple", "general", "reasoning", "code", "agent"}
return category if category in valid else "general"

The key details: max_tokens=10 because we only need one word. temperature=0 for deterministic routing. The system prompt is explicit about format — no preamble, just the category.


Each category maps to a model optimized for that task:

# Model routing table
ROUTE_TABLE = {
"simple": "glm-5.2",
"general": "glm-5.2",
"reasoning": "MiniMax-M3",
"code": "MiniMax-M3",
"agent": "kimi-k2.7",
}
def route_request(user_message: str, conversation_history: list) -> str:
"""Classify and route a request to the appropriate model."""
category = classify_request(user_message)
model = ROUTE_TABLE[category]
response = client.chat.completions.create(
model=model,
messages=conversation_history + [
{"role": "user", "content": user_message}
],
stream=True
)
# Stream the response back
full_response = ""
for chunk in response:
if chunk.choices[0].delta.content:
content = chunk.choices[0].delta.content
full_response += content
print(content, end="", flush=True)
return full_response

Notice that simple requests route back to GLM 5.2 — the same model that did the classification. For simple queries, the router overhead is effectively zero because the specialist is the same model and can reuse the warm connection.


The basic router works for most traffic, but production systems need a few refinements:

def route_request_production(
user_message: str,
conversation_history: list,
force_model: str = None
) -> tuple[str, str]:
"""Production router with overrides and fallback."""
# Allow explicit model override (for power users or testing)
if force_model:
model = force_model
category = "override"
else:
category = classify_request(user_message)
model = ROUTE_TABLE[category]
try:
response = client.chat.completions.create(
model=model,
messages=conversation_history + [
{"role": "user", "content": user_message}
]
)
return response.choices[0].message.content, category
except Exception:
# Fallback to GLM 5.2 if the specialist is unavailable
fallback = "glm-5.2"
response = client.chat.completions.create(
model=fallback,
messages=conversation_history + [
{"role": "user", "content": user_message}
]
)
return response.choices[0].message.content, f"{category}->fallback"

Three patterns worth noting:

  1. Force model — Let callers bypass routing when they know what they need.
  2. Fallback — If a specialist model is down, fall back to GLM 5.2. It handles everything reasonably well.
  3. Return the category — Log which route each request takes. You’ll need this data to tune the system.

Consider a workload of 1,000 requests with this distribution: 600 simple, 300 general, 70 reasoning, 30 code. Average 500 input tokens, 200 output tokens per request.

Single-model approach (everything on V3.2)

Section titled “Single-model approach (everything on V3.2)”
Avg latency
~4.5s
All 1000 reqs
V3.2 only

Every request waits for V3.2’s ~1.2s TTFT plus generation time at ~60 t/s. Simple questions get the same treatment as complex analysis.

Simple (600)
~1.2s (8B)
General (300)
~4.7s (V3.2)
Reasoning (70)
~9.0s (R1)
Code (30)
~3.5s (Coder)

The weighted average latency drops to approximately 2.7s — a 40% reduction. The 600 simple requests finish in ~1.2s instead of ~4.5s. That’s a 3.7x improvement for the majority of your traffic.

The 70 reasoning requests are slower individually (~9s vs ~4.5s) because R1 generates chain-of-thought tokens. But the quality on those specific requests is significantly better — R1 scores 50.2% on HLE versus V3.2’s 39.3%.

You get faster averages and better quality on the hard tail.


A customer support chatbot receives three types of requests:

  1. FAQ (60%) — “What are your business hours?” / “How do I reset my password?”
  2. Complex support (30%) — “I was charged twice for order #12345, can you investigate?”
  3. Technical issues (10%) — “Your API returns 500 when I send multipart form data with UTF-8 filenames”

All requests go to DeepSeek V3.2. FAQs get correct answers but with unnecessary latency. Technical issues get decent answers but miss edge cases that a code-specialized model would catch.

SUPPORT_ROUTES = {
"simple": "glm-5.2", # FAQ, greetings
"general": "glm-5.2", # Complex support
"reasoning": "glm-5.2", # Investigations
"code": "glm-5.2", # Technical issues
"agent": "kimi-k2.7", # Multi-step resolution
}

FAQs resolve quickly via GLM 5.2. Complex support issues get GLM 5.2’s full analytical capability. Technical problems also route to GLM 5.2, which understands the code context well. If a support issue requires looking up order data via API, it routes to Kimi K2.7 for tool-assisted resolution.

The classification step adds ~200ms. For the 60% of requests that drop from ~4.5s to ~1.2s, that’s an invisible cost.


Routing adds complexity. Skip it when:

  • All your requests are the same type. If you’re building a code editor, just use a single coding model like GLM 5.2. No routing needed.
  • You have fewer than 100 requests/day. The cost savings don’t justify the engineering overhead at low volume.
  • Latency doesn’t matter. For batch processing or async workloads, a single capable model is simpler.
  • Your classification accuracy is low. If the router misclassifies frequently, you get worse results than a single good model. Test the classifier on real traffic before deploying.

The sweet spot is high-volume applications with diverse request types — chatbots, API gateways, developer tools, and customer-facing products where response time directly affects user experience.


  1. Log your traffic. Before building a router, understand your request distribution. What percentage is simple? Complex? Code?
  2. Start with two tiers. A fast, lighter model for simple requests, and a stronger model like MiniMax M3 for everything that needs deep reasoning, code, or long context. Add specialists only when you have data showing they help.
  3. Measure classification accuracy. Sample 100 requests, manually label them, compare against the router’s output. Target >90% accuracy.
  4. Add fallback. Every specialist route should fall back to GLM 5.2 if the specialist is unavailable.
  5. Monitor per-route metrics. Track latency, cost, and quality per category. This tells you where to optimize next.

The routing pattern works with any OpenAI-compatible API. The code examples in this guide use the model ids we actually serve — GLM 5.2, MiniMax M3, and Kimi K2.7 from our Frontier pool (DeepSeek V4 Flash and MiMo v2.5 in the Core pool make great router models too); the throughput and latency comparisons cite ecosystem reference models like Llama and DeepSeek V3.2/R1 for context. If you’re building a platform that needs LLM access for your users, see how per-key plans work.

Sources: Artificial Analysis Leaderboard · DeepSeek V3.2 · HLE Leaderboard

What it takes to build your own LLM inference platform

If you’re building a SaaS that needs to give users access to LLMs, you have two options: build the infrastructure yourself, or use a platform that does it for you. Here’s what “build it yourself” actually looks like.

This isn’t theoretical. We built this. Here’s every component, what it does, and what alternatives exist.

Before you write a single line of code, you need access to models.

Self-host on your own hardware: Buy GPUs, rent datacenter space, run the models yourself. Full control, best unit economics at scale — but massive upfront cost and you’re limited to the models you can afford to deploy. Running DeepSeek V3.2 requires multiple high-end GPUs. Running dozens of models? You’d need a data center.

Rent infrastructure: Use GPU clouds like Vast.ai, AWS, Hetzner, CoreWeave, or Lambda. No hardware to buy, but you still manage deployments, scaling, and failover. Costs add up fast — a single H100 runs $2-4/hr.

Use an inference provider: Sign agreements with DeepInfra, Together.ai, Fireworks, etc. who already have the models deployed. Pay per token, no GPU management. But you depend on their availability, pricing, and terms. If they change prices or drop a model, you need a plan B.

Mix: Most serious platforms end up here. Own hardware for high-volume models where the unit economics justify it, rented GPUs for burst capacity, and provider agreements for the long tail of models nobody runs enough to self-host.

Self-hosting dozens of models on your own is economically unrealistic. The real question is where to draw the line between own infra, rented compute, and providers.

If you self-host or rent GPUs, you need software to serve the models:

  • vLLM — most popular, good throughput, active community
  • TGI (Text Generation Inference) — Hugging Face’s solution, solid for single-model deployments
  • TensorRT-LLM — NVIDIA’s optimized engine, best raw performance but harder to set up
  • SGLang — newer, fast, good for structured generation

You’ll also need to handle model weights, quantization, scaling across GPUs, and failover when a node goes down. This is a full-time ops job.

Your users shouldn’t hit the inference backend directly. You need a proxy that:

  • Translates between API formats (OpenAI, Anthropic)
  • Routes requests to the right model/provider
  • Injects authentication
  • Handles retries and failover
  • Strips provider headers so users don’t know your backend

Options:

  • Build from scratch with Express/Fastify + http-proxy-middleware
  • Use an open-source gateway: LiteLLM, Portkey, Kong AI Gateway, MLflow Gateway
  • Use a managed gateway: Helicone, Braintrust, Promptlayer

Each has trade-offs. Open-source gateways give you control but you manage the deployment. Managed gateways are easier but add latency and cost.

Two layers:

User auth (dashboard login)

  • Firebase Auth, Auth0, Clerk, Supabase Auth, or roll your own
  • Supports email, Google, GitHub, wallet signatures

API key auth (inference requests)

  • Generate API keys per user
  • Validate on every request before proxying
  • Store key metadata (plan, rate limits, owner)

This is where it gets interesting for platforms. You need per-key plans — each key with its own rate limits and usage tracking. Most auth solutions don’t do this out of the box. You’ll need a custom key management layer.

Per-key rate limiting with at least:

  • RPM (requests per minute)
  • TPM (tokens per minute)
  • Budget caps (dollar amount per time window)

This needs to be enforced at the proxy layer, before the request hits the inference backend. Otherwise a single user can exhaust your GPU allocation.

Options:

  • Redis-based counters (most common)
  • Token bucket algorithms
  • Proxy-level enforcement (some gateways include this)

If you’re using per-key plans, each key needs its own set of limits. Not one global limit — individual limits per key.

You need to know:

  • How many tokens each key consumed (input + output)
  • What model was used
  • Cost per request
  • Aggregate usage per user, per day, per billing period

For subscription billing:

  • Stripe for card payments
  • Budget windows (e.g., $X per 8-hour period)
  • Automatic key revocation when subscription expires

For pay-as-you-go:

  • Credit balance per user
  • Deduct per request based on token count × model price
  • Top-up flow (Stripe, crypto, etc.)

For crypto payments:

  • USDC on a supported chain
  • On-chain transaction verification
  • Wallet connector in the dashboard (wagmi, viem, etc.)

This is a significant amount of code. Usage tracking alone requires intercepting every response to count tokens, calculating cost based on the model’s pricing, and storing it per key.

Your users need a web UI to:

  • Create and manage API keys
  • View usage per key (tokens, requests, cost)
  • Subscribe to plans or top up credits
  • See available models and pricing

Tech stack typically:

  • React/Next.js/Vue frontend
  • REST API backend
  • Real-time usage updates

For platforms (your users creating keys for their users), you also need a management API — programmatic key creation, plan assignment, usage queries.

Models change. New ones come out weekly. You need:

  • A catalog of which models you serve
  • Pricing per model (input/output cost per token)
  • Sync mechanism to update prices when providers change them
  • Display names, categories, tags for the dashboard
  • Cache pricing metadata (some models support prompt caching discounts)

This is an ongoing operational burden, not a one-time setup.

Your users need:

  • API reference (endpoints, request/response formats)
  • SDK examples (Python, Node.js, at minimum)
  • Authentication guide
  • Billing/usage documentation
  • Quick start guide

This is easily 20-30 pages of documentation that needs to stay current.

  • Health checks on the inference backend
  • Status page for users
  • Alerting when latency spikes or errors increase
  • Logging (but not logging prompt content — privacy)
  • Graceful degradation when a model or provider is down
  • Privacy policy
  • Data handling documentation
  • GDPR compliance if you serve EU users
  • Decision: do you store prompts? (You shouldn’t)
  • SOC 2 / ISO 27001 if targeting enterprise

ComponentOngoing maintenance
Inference backendHigh — scaling, failover, model updates
API proxyMedium — format changes, new providers
Auth + key managementLow
Per-key rate limitingLow
Usage tracking + billingMedium — edge cases, reconciliation
DashboardMedium — new features, UX
Model catalogHigh — weekly model updates
DocumentationMedium — keep current
MonitoringLow
Privacy/complianceLow

Building is the easy part. The hard part is what breaks with real users:

  • A provider changes their API format without warning. Your proxy returns 500s for 2 hours until you notice.
  • A model gets deprecated. Your users’ hardcoded model IDs stop working overnight.
  • Token counting has an off-by-one bug. You’ve been undercharging for 3 weeks. Your margin is gone.
  • A user finds a way to exceed rate limits through concurrent requests. Your inference bill spikes 10x in one afternoon.
  • Stripe webhook fails silently. A user’s subscription expired but their API key still works. Free inference for a month.
  • You push a billing update and break the usage tracking. Three days of missing data. Users open tickets.

Each of these has happened to us. We fixed them. The question is whether you want to fix them yourself, with your users waiting, or use a platform that already has.

You use an inference platform that already has all of this, create API keys for your users, and ship your product this week.


We built all of the above so you don’t have to. See how per-key plans work.