Guides / forecast accuracy
What is forecast accuracy?
Forecast accuracy is two measurements taken together: how far the middle estimate sat from what actually happened, and whether reality landed inside the upper band. Either one on its own can flatter a useless forecast. Culpa, a local-first LLM cost, margin, and forecast ledger, forecasts every period and grades it once the period closes, so the record is scored rather than remembered.
Why this happens
An ungraded forecast is a guess with a decimal point, and grading it on one axis is barely better. Score only the middle estimate and a forecast that happens to centre well looks excellent while its band admits anything. Score only the band and a wide enough range is never wrong and never useful. The pair is what carries information: a tight band that reality keeps landing inside means the model understands your spend, and a wide band nobody ever escapes means it has learned to say nothing safely.
What this usually looks like
- Forecasts get made and never compared against what happened.
- Accuracy is discussed as a single percentage with no mention of the band.
- Your bands are wide enough that reality has never fallen outside one.
- Nobody can say whether last quarter's forecasts were any good.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Mistakes that cost the most
| Mistake | Why it hurts | Do instead |
|---|---|---|
| Grading a forecast on its middle estimate alone. | It says nothing about the range you actually planned against, which is what bounds your risk. | Record the error on the centre and whether the outcome fell inside the band, every period. |
| Widening the band until nothing ever falls outside it. | A forecast that can't be wrong can't inform a decision either. | Track band width alongside hit rate. A band that never binds is a band that never helped. |
| Not storing the forecast that was made. | You can't grade what you didn't keep, so the accuracy question becomes unanswerable forever. | Persist every forecast at the time it's made, then score it when the period closes. |
Run this check tonight
- Find last month's forecast. If it wasn't written down, that's the first finding.
- Compute the middle estimate's error against the actual as a percentage.
- Check whether the actual landed inside the upper band.
- Divide the band width by the middle estimate. A ratio above two admits almost anything.
- Do all four for the last six periods, and look at the trend rather than any single month.
Four months graded on both axes
Illustrative example
Four closed monthly periods, each with a forecast median, an upper band and an actual. Error is the absolute difference from the actual, as a percentage of the actual. Figures are modelled.
Month 4 scores best on the centre and tells you least, because a band more than double the median would have absorbed almost any outcome. Month 3 is the only period that actually failed, and only the band shows it.
Every number, with its confidence and source
| Figure | What it means | Confidence | Source |
|---|---|---|---|
| 3.23% to 31.03% | modelled median forecast error across four graded periods | estimated | Each endpoint is the absolute difference between the modelled median and actual, over the actual, from the teardown arithmetic. A range because the four periods are modelled. |
| 2.3x | modelled band width as a multiple of the median in the period that scored best on error | calculated | $700 divided by $300, from the month 4 figures in the teardown. |
What a generic answer can’t know
Nobody outside your own history can tell you whether your forecasts have been any good, because grading one needs the forecast as it was made and the outcome that followed, held together over time. Culpa keeps both on your infrastructure, where your prompts and responses stay, and counts the calls to run your plan.
Questions founders ask next
How should an LLM spend forecast be graded?
On two numbers. The percentage error between the middle estimate and the actual, and whether the actual fell inside the upper band. Reporting either alone lets a forecast that centred by luck, or one that hedged into uselessness, look good.
What's a good median error on LLM spend?
There's no defensible figure to quote across products, because volatility differs enormously with traffic shape. The useful target is your own trend, plus a band that binds often enough to mean something when reality stays inside it.
Why does the band matter as much as the estimate?
Because the band is what you plan against. A middle estimate tells you the likely case and the band tells you the case you have to survive, so a forecast graded without it says nothing about whether your budget was ever safe.
On your infrastructure
Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.
How Culpa works
Find the culprit. Not just the total.
Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.
Your prompts stay local.
Culpa runs on your own infrastructure. What you send to a model reaches us at no point.
Every dollar has a name.
Follow any charge to the conversation, the user, the feature and the customer behind it.
See the bill before it lands.
Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.
Three steps to your first answer.
Change one base URL.
Or drop in the Python or TypeScript library.
Find your most expensive conversation.
In the first session, not the first week.
Cost your next feature before you ship it.
Why the bill went up
Example dashboardCalls traced
418,209
across 3 projects
Spend this week
$378.41
+ $182 vs last week
Failed calls
312
74% retried, and you paid for all of them
+ $182 this week traced to one culprit
Spend over 14 days
Most expensive users
Next week forecast
Graded against reality. Accuracy shown as results land.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Keep reading
Sources: OpenAI API pricing. Last reviewed 2026-08-02. Plain text version.