Skip to content

Flat Rate vs Per-Token — the break-even curve and the savings factor, updated monthly

Living report · updated monthly, with each Pareto edition, and on every price change · last updated

Every conversation about inference cost lands on the same question: at what point does a flat monthly block beat paying per token? The answer is a curve, not a number. This report draws it, computes the break-even volume for every pool from the vendors’ own published per-token list prices, and states the savings factor honestly — including the ceiling a single key will not go past.

Two inputs, both live: our flat prices (the current cheapest block of each pool; billed annually, Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo) and the per-token list prices tracked in the LLM Price Tracker. Every figure below is recomputed at build time from those sources; nothing is hand-typed.

Per-token cost is a straight line through the origin: twice the tokens, twice the bill. A flat block is a horizontal line: the invoice is decided when you subscribe. The two cross at the break-even volume. To the left, per-token is cheaper. To the right, the flat block is, and the gap widens with every extra token.

Per-token, list price of the comparable alternatives (band: cheapest to priciest)Flat Core block, from $22.00/mo
0100M200M300M400M500M tokens per month $0$125$250 Per-token, priciest comparable alternative: $0.48 per million tokens blended Per-token, cheapest comparable alternative: $0.17 per million tokens blended Flat Core block: $22.00 per month, unlimited tokens during the reserved hours Break-even against the priciest alternative: about 46M tokens per month Break-even against the cheapest alternative: about 131M tokens per month break-even ≈ 46M–131M tokens/month per-token: ≈ $75.60–$216.00 at 450M flat block: $22.00, any volume

The chart shows the Core pool at its current monthly price; the other pools follow the same shape at their own prices (table below).

Break-even per pool

One row per pool, against the published per-token list prices of the comparable alternatives, as a range from the cheapest to the priciest. Blended rate: 80% input and 20% output, no cache discount. Flat price is the pool's current cheapest block, billed monthly; the last column is the same break-even at the annual-billed price. Prices as of edition 2026-09; current pricing is on the pools page.

PoolPer-token list price of the alternativesBreak-even volumePer day (30-day month)Billed annually
Flagship, from $199.00/mo ≈ $2.80–$5.40 / M tokens ≈ 37M–71M tokens/month ≈ 1.2M–2.4M ≈ 31M–60M tokens/month (from $169.15/mo)
Frontier, from $71.00/mo ≈ $0.48–$2.00 / M tokens ≈ 36M–148M tokens/month ≈ 1.2M–4.9M ≈ 30M–126M tokens/month (from $60.35/mo)
Core, from $22.00/mo ≈ $0.17–$0.48 / M tokens ≈ 46M–131M tokens/month ≈ 1.5M–4.4M ≈ 39M–111M tokens/month (from $18.70/mo)

The savings factor

Per-token cost of the same tokens divided by the flat monthly price, per pool, as a range across the comparable alternatives. Three usage profiles, all at or above break-even: below roughly 130–150M tokens a month the per-token line can be cheaper — see the break-even row above. The heavy column assumes high cache reuse or more than one subscription; it is not a level of service to expect — see "What unlimited and fair use mean here" below.

PoolBreak-even and above
160M tokens/mo
Daily driver
230M tokens/mo
Heavy
300M tokens/mo
Flagship, from $199.00/mo ≈ 2.3×–4.3×≈ 3.2×–6.2×≈ 4.2×–8.1×
Frontier, from $71.00/mo ≈ 1.1×–4.5×≈ 1.6×–6.5×≈ 2.0×–8.5×
Core, from $22.00/mo ≈ 1.2×–3.5×≈ 1.8×–5.0×≈ 2.3×–6.5×

No SLA. Throughput, latency and the savings factor are not guaranteed: they depend on your workload mix and on the capacity available in the pool at the time. Plan with the daily-driver column; treat the heavy column as an upper bound, not as an expectation.

For scale: a single coding-agent task typically burns 300–500K tokens, because the agent re-sends its growing context on every tool call, and we measured a coding agent at roughly 2M tokens per hour. One to two million tokens a day is a handful of tasks, not a heavy day. Most people who run an agent daily are past break-even in the first week of the month.

What “unlimited” and fair use mean here

Section titled “What “unlimited” and fair use mean here”

Read the two tables with the right expectations, because a flat block is a different product from a meter:

  • Unlimited means tokens. Nothing is metered or billed per token during your reserved hours, and there are no overage charges. That is what makes the factor grow with your volume instead of the bill.
  • One request at a time per key. Throughput comes from concurrency, and concurrency comes from subscriptions: fold several into a combined key and parallel capacity stacks where their hours overlap.
  • Pools are shared, so fair use applies. When one key’s usage over a billing period runs far above what a single subscriber’s workload normally represents, the platform moderates that key’s speed — depending on how much capacity the pool has available at the time — so the pool keeps working for everyone. It slows; it does not cut off, and it never bills. The calibration is internal and not published; the rule itself is in our Terms, section 1.2.

Why we do not publish a sharper number. What a request costs to serve is not a token count: it depends on how much of the input is fresh versus served from cache, on request size, on output length, and on how loaded the pool is at that moment. The same million tokens can cost several times more or less depending on that mix. Moderating a key is also only one of several levers — capacity, pricing and moderation are balanced together, and that balance is the product. A single published threshold would be wrong for most workloads and right only for whoever tuned against it, so we publish the rule, not a figure.

So the columns read like this:

  • Break-even and above is the volume from which the flat block wins against every comparable alternative in every pool. Below it, per-token still competes: if your month sits under the break-even row above, pay per token.
  • Daily driver is what an active developer or a busy agent actually reaches on one key: several times what the same month would cost per token. This is the column to plan with.
  • Heavy is what one key reaches with high cache reuse, or what you buy outright with more than one subscription. It is an upper bound, not a level of service you should expect: moderation applies under fair use as availability tightens, so treat anything above the daily-driver factor as upside — or buy it with more subscriptions, each with its own curve.
  • There is no SLA. We do not guarantee throughput, latency or a savings factor; the figures here are arithmetic on published prices and illustrative volumes, and the Terms govern.

That is the whole model in one sentence: flat rate wins by a factor that grows with your volume, up to what one key can push through its reserved block.

The curve is symmetric, so it also says when not to subscribe:

  • Sporadic use. A weekly batch job, an internal tool a few people touch once a day, a prototype you poke at on weekends. Below the break-even volume, per-token is cheaper and you should pay per token.
  • Usage concentrated outside your block. A block is a daily 8-hour window. If your traffic is spread evenly around the clock, you either need more blocks or you are paying for hours you do not use. Combined keys cover this: coverage windows add up across subscriptions.

If someone asks “is the subscription worth it for us?”, the answer is a volume, not an opinion:

  1. Estimate tokens per month per person or per agent. Count context re-sends; they dominate.
  2. Compare with the break-even column for the pool you would use.
  3. If you are past it, the flat block is cheaper by the factor in the second table, capped at what one key can push through. If you are not, pay per token.

Is the savings factor guaranteed? No. It depends on your volume and on the throughput one key delivers. The daily-driver column is what a typical active developer sees; the heavy column is an upper bound, reached with high cache reuse or with more than one subscription, and moderated under fair use as availability tightens. Subscriptions are unlimited in tokens, not in throughput.

What if I need more than one key can deliver? Take out additional subscriptions and fold them into a combined key. Where their hours overlap, parallel capacity stacks; where they do not, coverage extends. Each subscription is its own flat block with its own curve.

Is there an SLA? No. We do not guarantee throughput, latency or any savings factor. Fair-use moderation is not a service failure and does not give rise to refunds or credits; the Terms govern. The figures in this report are arithmetic on published prices and illustrative volumes.

Where are the current prices? Always on the pools page. This report recomputes from them on every build; the page is the source of truth.

Each pool is compared against the published per-token list prices of the comparable alternatives for that class of model, shown as a range from the cheapest to the priciest. Per-token prices are vendor list prices for the variants Artificial Analysis evaluates (peak list price where a vendor publishes peak/off-peak); open-weight models are often cheaper on aggregators, so the per-token side is conservative in the vendor’s favour. The 80/20 input/output mix is typical of agent workloads; a chat-heavy mix has more output and a higher blended rate. Usage profiles are illustrative volumes, not measurements of any subscriber.

  • 2026-09-17 — First edition: break-even volume and savings factor per pool, recomputed from live prices and the monthly Pareto edition.