Flat Rate vs Per-Token — the break-even curve and the savings factor, updated monthly
Every conversation about inference cost lands on the same question: at what point does a flat monthly block beat paying per token? The answer is a curve, not a number. This report draws it, computes the break-even volume for every pool from the vendors’ own published per-token list prices, and states the savings factor honestly — including the ceiling a single key will not go past.
Two inputs, both live: our flat prices (the current cheapest block of each pool; billed annually, Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo) and the per-token list prices tracked in the LLM Price Tracker. Every figure below is recomputed at build time from those sources; nothing is hand-typed.
The curve
Section titled “The curve”Per-token cost is a straight line through the origin: twice the tokens, twice the bill. A flat block is a horizontal line: the invoice is decided when you subscribe. The two cross at the break-even volume. To the left, per-token is cheaper. To the right, the flat block is, and the gap widens with every extra token.
The chart shows the Core pool at its current monthly price; the other pools follow the same shape at their own prices (table below).
Break-even per pool
One row per pool, against the published per-token list prices of the comparable alternatives, as a range from the cheapest to the priciest. Blended rate: 80% input and 20% output, no cache discount. Flat price is the pool's current cheapest block, billed monthly; the last column is the same break-even at the annual-billed price. Prices as of edition 2026-09; current pricing is on the pools page.
| Pool | Per-token list price of the alternatives | Break-even volume | Per day (30-day month) | Billed annually |
|---|---|---|---|---|
| Flagship, from $199.00/mo | ≈ $2.80–$5.40 / M tokens | ≈ 37M–71M tokens/month | ≈ 1.2M–2.4M | ≈ 31M–60M tokens/month (from $169.15/mo) |
| Frontier, from $71.00/mo | ≈ $0.48–$2.00 / M tokens | ≈ 36M–148M tokens/month | ≈ 1.2M–4.9M | ≈ 30M–126M tokens/month (from $60.35/mo) |
| Core, from $22.00/mo | ≈ $0.17–$0.48 / M tokens | ≈ 46M–131M tokens/month | ≈ 1.5M–4.4M | ≈ 39M–111M tokens/month (from $18.70/mo) |
The savings factor
Per-token cost of the same tokens divided by the flat monthly price, per pool, as a range across the comparable alternatives. Three usage profiles, all at or above break-even: below roughly 130–150M tokens a month the per-token line can be cheaper — see the break-even row above. The heavy column assumes high cache reuse or more than one subscription; it is not a level of service to expect — see "What unlimited and fair use mean here" below.
| Pool | Break-even and above 160M tokens/mo | Daily driver 230M tokens/mo | Heavy 300M tokens/mo |
|---|---|---|---|
| Flagship, from $199.00/mo | ≈ 2.3×–4.3× | ≈ 3.2×–6.2× | ≈ 4.2×–8.1× |
| Frontier, from $71.00/mo | ≈ 1.1×–4.5× | ≈ 1.6×–6.5× | ≈ 2.0×–8.5× |
| Core, from $22.00/mo | ≈ 1.2×–3.5× | ≈ 1.8×–5.0× | ≈ 2.3×–6.5× |
No SLA. Throughput, latency and the savings factor are not guaranteed: they depend on your workload mix and on the capacity available in the pool at the time. Plan with the daily-driver column; treat the heavy column as an upper bound, not as an expectation.
For scale: a single coding-agent task typically burns 300–500K tokens, because the agent re-sends its growing context on every tool call, and we measured a coding agent at roughly 2M tokens per hour. One to two million tokens a day is a handful of tasks, not a heavy day. Most people who run an agent daily are past break-even in the first week of the month.
What “unlimited” and fair use mean here
Section titled “What “unlimited” and fair use mean here”Read the two tables with the right expectations, because a flat block is a different product from a meter:
- Unlimited means tokens. Nothing is metered or billed per token during your reserved hours, and there are no overage charges. That is what makes the factor grow with your volume instead of the bill.
- One request at a time per key. Throughput comes from concurrency, and concurrency comes from subscriptions: fold several into a combined key and parallel capacity stacks where their hours overlap.
- Pools are shared, so fair use applies. When one key’s usage over a billing period runs far above what a single subscriber’s workload normally represents, the platform moderates that key’s speed — depending on how much capacity the pool has available at the time — so the pool keeps working for everyone. It slows; it does not cut off, and it never bills. The calibration is internal and not published; the rule itself is in our Terms, section 1.2.
Why we do not publish a sharper number. What a request costs to serve is not a token count: it depends on how much of the input is fresh versus served from cache, on request size, on output length, and on how loaded the pool is at that moment. The same million tokens can cost several times more or less depending on that mix. Moderating a key is also only one of several levers — capacity, pricing and moderation are balanced together, and that balance is the product. A single published threshold would be wrong for most workloads and right only for whoever tuned against it, so we publish the rule, not a figure.
So the columns read like this:
- Break-even and above is the volume from which the flat block wins against every comparable alternative in every pool. Below it, per-token still competes: if your month sits under the break-even row above, pay per token.
- Daily driver is what an active developer or a busy agent actually reaches on one key: several times what the same month would cost per token. This is the column to plan with.
- Heavy is what one key reaches with high cache reuse, or what you buy outright with more than one subscription. It is an upper bound, not a level of service you should expect: moderation applies under fair use as availability tightens, so treat anything above the daily-driver factor as upside — or buy it with more subscriptions, each with its own curve.
- There is no SLA. We do not guarantee throughput, latency or a savings factor; the figures here are arithmetic on published prices and illustrative volumes, and the Terms govern.
That is the whole model in one sentence: flat rate wins by a factor that grows with your volume, up to what one key can push through its reserved block.
When per-token wins
Section titled “When per-token wins”The curve is symmetric, so it also says when not to subscribe:
- Sporadic use. A weekly batch job, an internal tool a few people touch once a day, a prototype you poke at on weekends. Below the break-even volume, per-token is cheaper and you should pay per token.
- Usage concentrated outside your block. A block is a daily 8-hour window. If your traffic is spread evenly around the clock, you either need more blocks or you are paying for hours you do not use. Combined keys cover this: coverage windows add up across subscriptions.
How to use this
Section titled “How to use this”If someone asks “is the subscription worth it for us?”, the answer is a volume, not an opinion:
- Estimate tokens per month per person or per agent. Count context re-sends; they dominate.
- Compare with the break-even column for the pool you would use.
- If you are past it, the flat block is cheaper by the factor in the second table, capped at what one key can push through. If you are not, pay per token.
Common questions
Section titled “Common questions”Is the savings factor guaranteed? No. It depends on your volume and on the throughput one key delivers. The daily-driver column is what a typical active developer sees; the heavy column is an upper bound, reached with high cache reuse or with more than one subscription, and moderated under fair use as availability tightens. Subscriptions are unlimited in tokens, not in throughput.
What if I need more than one key can deliver? Take out additional subscriptions and fold them into a combined key. Where their hours overlap, parallel capacity stacks; where they do not, coverage extends. Each subscription is its own flat block with its own curve.
Is there an SLA? No. We do not guarantee throughput, latency or any savings factor. Fair-use moderation is not a service failure and does not give rise to refunds or credits; the Terms govern. The figures in this report are arithmetic on published prices and illustrative volumes.
Where are the current prices? Always on the pools page. This report recomputes from them on every build; the page is the source of truth.
Caveats
Section titled “Caveats”Each pool is compared against the published per-token list prices of the comparable alternatives for that class of model, shown as a range from the cheapest to the priciest. Per-token prices are vendor list prices for the variants Artificial Analysis evaluates (peak list price where a vendor publishes peak/off-peak); open-weight models are often cheaper on aggregators, so the per-token side is conservative in the vendor’s favour. The 80/20 input/output mix is typical of agent workloads; a chat-heavy mix has more output and a higher blended rate. Usage profiles are illustrative volumes, not measurements of any subscriber.
Changelog
Section titled “Changelog”- 2026-09-17 — First edition: break-even volume and savings factor per pool, recomputed from live prices and the monthly Pareto edition.