Guides / per customer llm cost forecast

Forecasting LLM cost per customer, not just in total

A per-customer LLM cost forecast projects each account's spend from its own history rather than projecting one company-wide total. The aggregate can land perfectly while the customers inside it move in opposite directions. Culpa, a local-first LLM cost, margin, and forecast ledger, forecasts at the level it attributes, so a projection names an account.

Why this happens

Almost every cost forecast in this category is a single line for the whole company, and a single line has a specific failure mode: it can be accurate and useless at the same time. If one customer doubles while another halves, the total is unchanged and both facts that matter are invisible. The reason forecasts are aggregate is rarely a modelling choice. It's that the underlying data was aggregated first, so nothing finer remains to project. Datadog's Cloud Cost Management forecasts spend built from an ingested vendor invoice, which is the coarsest object in the stack and contains no customer at all. Dynatrace says it can predict cost increases, which tells you a number is trending up rather than where any account lands. Across the twelve LLM-native tools checked on 2026-08-03, none publishes a spend forecast of any kind. So the forecast ceiling sits at or above the attribution ceiling in every case, which follows necessarily: you can't project a thing you never recorded separately.

What this usually looks like

  • Your forecast was accurate last month and two accounts still surprised you.
  • A customer's usage tripled and the company total absorbed it.
  • You can't tell a renewal conversation what that account will cost next quarter.
  • Growth and churn cancel out in the total and nobody sees either.
  • A forecast exists and no account manager has ever been shown one.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run the free Cost Leak ScanStart 14-day trial

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Forecasting one line for the whole company.Opposing movements cancel, so the total can be right while every component is wrong.Forecast each account separately and let the total be the sum, not the input.
Projecting from an ingested vendor invoice.An invoice has no customer, feature or prompt version in it, so no breakdown can come out of it.Project from your own per-call ledger, where the attribution already exists.
Treating a cost-increase alert as a per-account forecast.It tells you something rose, after it rose. It doesn't tell you which account or where it lands.Keep the alert and add a projection per account, with a range.
Forecasting spend without the revenue beside it.A customer growing fast is good news or bad news depending entirely on what they pay you.Forecast margin per account, which means forecasting cost and holding revenue in the same place.

Run this check tonight

  1. Take your three largest accounts and write down what each will cost next month.
  2. Check whether any system could have produced those three numbers for you.
  3. Find an account whose usage moved sharply and see whether the company total shows it.
  4. Ask what your fastest-growing account's margin looks like at next quarter's volume.

An aggregate forecast, exactly right and completely useless

Three modelled accounts across two months, priced on Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book effective 2026-07-02. The company total is flat across both months, so an aggregate forecast projecting no change would score perfectly. Account volumes are modelled.

month 1: A $400.00, B $300.00, C $200.00, total $900.00
month 2: A $700.00, B $150.00, C $50.00, total $900.00
the aggregate forecast said $900.00 and was exactly right
account A grew 75%, account C fell 75%, and neither appears in the total

A forecast that scores perfectly and misses a customer tripling toward unprofitability is the clearest argument for forecasting where you attribute. The accuracy was real. The usefulness was zero.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
$900.00modelled company total, identical across two months while one account grew 75% and another fell 75%calculatedThree modelled accounts priced on Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book effective 2026-07-02. Month 1: $400.00 + $300.00 + $200.00 = $900.00. Month 2: $700.00 + $150.00 + $50.00 = $900.00. Account volumes are modelled and chosen so the total is flat, which is the entire point of the figure.

What a generic answer can’t know

A forecast can only be as granular as the data underneath it, which is why the forecast ceiling and the attribution ceiling are the same ceiling in practice. A platform that stores cost per agent can forecast per agent at best. One that ingests a vendor invoice can forecast that invoice and nothing inside it. Culpa prices every call individually and tags it with the customer, the feature and the conversation, so a forecast can be run at any of those levels rather than only at the top, and each one is published as a range because a single number hides how much that account's own months disagree. Every forecast is persisted and scored against the actual when the month closes, which is what makes accuracy a number rather than a claim.

Questions founders ask next

Why forecast per customer rather than in total?

Because opposing movements cancel in a total. One account tripling while another collapses leaves the company figure unchanged, so an aggregate forecast can score perfectly while missing both events. The total is worth having as a sum of the parts rather than as the thing you project.

Why don't other tools forecast per customer?

Mostly because their data was aggregated before it reached the forecast. Datadog's Cloud Cost Management projects spend built from an ingested vendor invoice, which contains no customer. Across the twelve LLM-native tools checked on 2026-08-03, none publishes a spend forecast at all.

How much history does a per-account forecast need?

More than most retention tiers hold, which is the practical catch. Free tiers in this category run from 24 hours to 90 days, and a seasonal read on one account needs a year. Keeping the ledger on storage you own is what makes the lookback a disk decision rather than a plan tier.

What makes a per-account forecast trustworthy?

Publishing it as a range rather than a point, and scoring it afterwards. A projection nobody compares to the actual is an opinion. Culpa persists every forecast so accuracy can be measured against what happened rather than asserted in a sales page.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Keep reading


Sources: Anthropic pricing, Datadog LLM Observability. Last reviewed 2026-08-03, rates effective 2026-07-02. Plain text version.