Guides / llm cost forecast

LLM cost forecast calculator

An LLM cost forecast projects next month's spend from your own history, and it has to be a range because a single number hides how much your months disagree. This calculator prices three months you enter and projects the next one. Culpa, a local-first LLM cost, margin, and forecast ledger, forecasts from your full ledger and scores every forecast afterwards.

Computed in your browser, in exact decimal

Last month, priced$2,590.00
Your observed growth+25.0% then +23.3%

Next month, slowest observed growth$3,194.33
Midpoint of the band$3,215.92
Next month, fastest observed growth$3,237.50

Band width, against last month1.6%

These bounds are the range your own growth already spanned, not a confidence interval. Two growth steps can't support a p90 or a standard deviation, and publishing one would be inventing precision you don't have. A wide band here is information: it means recent months disagreed with each other.

Rates come from Culpa's price book, effective 2026-07-02. Every figure is computed in exact decimal rather than floating point, which is why totals here match an invoice to the cent. Cost is calculated, not provider-reported: it prices the tokens you enter at a published rate.

Why this happens

Most LLM spend forecasts are last month plus a percentage somebody chose, which is a budget wearing a forecast's clothes. A real projection has to come from your own history and it has to carry its uncertainty on its face. This calculator does the smallest honest version of that. It reads the growth steps between the months you type, then projects next month twice: once at the slowest growth you actually saw and once at the fastest. The gap between those two is the answer, and the width of the gap is the part worth reading. A narrow band means your recent months agreed with each other. A wide band means they didn't, and no amount of arithmetic can turn disagreement into confidence. What this deliberately doesn't do is publish a p90 or a confidence interval, because two growth steps can't support either and a statistic invented from two points is worse than no statistic at all.

What this usually looks like

  • Your forecast is last month's bill with a percentage added, and nobody remembers who picked the percentage.
  • Finance asks for one number and the honest answer has a range around it.
  • A month came in 40% over plan and the plan had never expressed a worst case.
  • You can see last month's spend and nothing about next month's.
  • Nobody checks afterwards whether the last forecast was any good.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run the free Cost Leak ScanStart 14-day trial

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Publishing a forecast as a single number.It hides the disagreement between your own months, which is the most useful thing the data holds.Publish a band, and let its width carry the uncertainty.
Calling a growth assumption a forecast.Last month plus 20% tells you about the 20%, which you chose, rather than about your traffic.Derive the bounds from growth that actually happened, then say which months produced them.
Quoting a p90 from three months of data.Two growth steps support no percentile at all. The decimal point implies precision that isn't there.Name the range you observed and say plainly how many observations produced it.
Forecasting tokens and forgetting the mix.Output usually costs several times input, so the same token growth costs differently depending on the split.Project the mix as well as the volume, and price each bucket at its own rate.
Never scoring the forecast afterwards.A forecast nobody checks is an opinion. Accuracy only exists if last month's projection was kept.Persist every forecast and compare it to the actual when the month closes.

Run this check tonight

  1. Write down next month's spend as a range before you look at anything.
  2. Find the last three months of actual spend and compute the two growth steps.
  3. Check whether your band from those steps contains the number you first wrote.
  4. Ask where last month's forecast is recorded, and whether anyone scored it.
  5. Decide which number you would defend to finance: the midpoint or the top of the band.

Two months that grow the same amount and forecast very differently

Take two products that both ended last month at 200M tokens, priced on Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book, effective 2026-07-02, at a 10% output share. One grew smoothly, the other lurched. Volumes are modelled and the arithmetic is the calculator's own.

both products priced last month at 200M tokens, 10% output share: $280.00
steady product: 165M, 182M, 200M tokens, growth steps of +10.3% then +9.9%
lurching product: 100M, 190M, 200M tokens, growth steps of +90.0% then +5.3%
steady forecast band: $307.69 to $308.85, a spread of $1.16
lurching forecast band: $294.74 to $532.00, a spread of $237.26

Both products spent the same last month and both are growing. One can be planned against and the other can't, and a single-number forecast would have reported them identically. The band width is the finding, not the midpoint.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
$294.74 to $532.00forecast band for a modelled product whose growth lurched from +90.0% to +5.3%estimatedComputed by this page's calculator from months of 100M, 190M and 200M tokens at a 10% output share, priced on Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book effective 2026-07-02. Observed growth steps are +90.0% then +5.3%, so the bounds are last month's 200M grown by each. Token volumes are modelled. Published as a range because it's an estimate, and because two growth steps support a range and nothing narrower.
$307.69 to $308.85forecast band for a modelled product at the same last-month volume, growing steadilyestimatedSame method and same rates, from months of 165M, 182M and 200M tokens at a 10% output share, giving growth steps of +10.3% then +9.9%. Both products ended last month at 200M tokens and both priced at $280.00. The steady band spans $1.16 and the lurching one spans $237.26, a factor of 204, which is the point of the pair.

What a generic answer can’t know

This calculator is deliberately the weak version, and saying so is the point. It sees three numbers you typed, so it can only read two growth steps, and two steps support a range and nothing more. What it does show is the shape of the honest answer: bounds drawn from your own history rather than a percentage somebody picked. Culpa runs the same idea over the whole ledger. It prices every call from a versioned price book in exact decimal rather than asking you to type a total, it projects from your full history rather than three points, it forecasts per customer and per feature rather than in aggregate, and it persists every forecast so that when the month closes it can score how wrong it was. A forecast nobody keeps is an opinion. Accuracy is only a number if last month's projection still exists.

Questions founders ask next

How do I forecast LLM costs?

Take your own recent months, measure the growth between them, and project the next month at both the slowest and fastest growth you observed. Publish the result as a range. Anything that reports a single number is hiding how much your months disagreed, and that disagreement is the most useful part.

Why won't this show a p90 or a confidence interval?

Because three months give two growth steps, and two observations support neither. A percentile computed from two points is a decimal point pretending to be evidence. The band here is exactly what it says: the range your own growth already spanned.

What does a wide band mean?

That your recent months disagreed with each other, which is information rather than a defect in the arithmetic. A product growing 90% then 5% genuinely can't be planned as tightly as one growing 10% twice, and a single-number forecast would have hidden that difference completely.

Is this how Culpa forecasts?

It's the same principle at a much smaller scale. Culpa projects from your full ledger rather than three typed numbers, works per customer and per feature rather than in aggregate, and persists every forecast so it can be scored against the actual once the month closes.

Why does the output share matter so much?

Because output tokens usually cost several times input tokens. On Claude Haiku 4.5 the published rates are $1.00 and $5.00 per million, so moving the output share from 10% to 20% raises the bill on identical token growth. Forecasting volume without the mix forecasts the wrong thing.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Keep reading


Sources: Anthropic pricing. Last reviewed 2026-08-02, rates effective 2026-07-02. Plain text version.