Guides / annual ai budget

Putting an AI line in next year's budget

An annual AI budget multiplies a run rate by twelve, and that arithmetic assumes both your usage and your rates hold still. Neither does, and rate changes arrive on published dates you can look up. Culpa, a local-first LLM cost, margin, and forecast ledger, forecasts from your own priced history against a price book versioned by date.

Why this happens

A budget is different from a forecast, and the difference is the reason this goes wrong. A forecast is a best guess you revise. A budget is a number you commit to and get measured against, so being wrong has consequences a forecast never carries. Most annual AI lines are built the same way: take the current month, multiply by twelve, add a round percentage for growth. That method has two blind spots. It treats a growth rate as a single number when it's a range, and it treats today's rates as next year's, which they aren't, because providers publish dated changes ahead of time. Claude Sonnet 5's introductory pricing ends on 2026-09-01 and rises 50%, and that date is public now. A budget built this month that runs through it lands already wrong by an amount you could have looked up. The fix is to budget a range rather than a point, price each month at the rate in effect for that month, and name what you would cut if the top of the range arrives.

What this usually looks like

  • Your annual number is this month times twelve plus a round percentage.
  • Nobody checked whether a published rate change lands inside the budget year.
  • The budget is a single figure with no range around it.
  • Growth was assumed flat because the alternative was arguing about it.
  • No agreed action exists for the case where spend runs above plan.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run the free Cost Leak ScanStart 14-day trial

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Multiplying the current month by twelve.It fixes both usage and rates, and neither holds for a year.Build month by month, each at the rate in effect for that month.
Budgeting a single number.A point estimate is wrong by construction and gives nobody a trigger to act on.Commit a range, and say in advance what happens at the top of it.
Ignoring published dated rate changes.They're knowable today, so missing one is an avoidable error rather than a forecasting risk.List every rate change with a date inside the budget year before you start.
Budgeting one AI line for the whole company.It can't be defended or cut, because nobody can see which product or team it belongs to.Split the line the way the business is run, so an overrun has an owner.

Run this check tonight

  1. Check which of your models have a published rate change dated inside your budget year.
  2. Rebuild the annual number month by month at the correct rate for each month.
  3. Put a range around it from your slowest and fastest recent growth, not a single figure.
  4. Write down what you would cut at the top of the range, before you need to.

Twelve months across a 50% rate rise

Illustrative example

A modelled workload of 10,000M input and 1,000M output tokens a month on Claude Sonnet 5, held flat all year so the rate change is the only thing moving. Sonnet 5's introductory rates are $2.00 and $10.00 per million, and the price book carries a second row effective 2026-09-01 at $3.00 and $15.00. The budget year starts in May, putting four months before the change and eight after. Token volumes and the budget year are modelled, both rates and the date are published.

at introductory rates: 10,000M x $2.00/M + 1,000M x $10.00/M = $20,000 + $10,000 = $30,000 a month
from 2026-09-01: 10,000M x $3.00/M + 1,000M x $15.00/M = $30,000 + $15,000 = $45,000 a month
run rate times twelve: 12 x $30,000 = $360,000
month by month: (4 x $30,000) + (8 x $45,000) = $120,000 + $360,000 = $480,000
the budget is $120,000 light, or 33.3% under, on zero growth

Usage never moved in this example. The entire $120,000 comes from one published date that anyone could have looked up before the budget was signed, and the run-rate method has no place to put it.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
$120,000, or 33.3%modelled annual shortfall from budgeting at a run rate across a published rate changecalculated10,000M input and 1,000M output tokens a month on Claude Sonnet 5 costs $20,000 + $10,000 = $30,000 at the introductory rates of $2.00 and $10.00 per million, and $30,000 + $15,000 = $45,000 at the $3.00 and $15.00 rates the price book carries effective 2026-09-01. Twelve months at the run rate is $360,000, against (4 x $30,000) + (8 x $45,000) = $480,000 built month by month, a shortfall of $120,000 or 33.3%. Both rates and the effective date are published. Token volumes and the budget year are modelled.

What a generic answer can’t know

Half of this is public and half of it belongs to you. The rate changes are published, dated, and the same for everybody, so missing one is an avoidable error rather than an act of forecasting. Your usage curve is the other half, and nothing outside your own history predicts it. Culpa forecasts from your own priced calls, deterministically and as a range rather than a single number, against a price book versioned by effective date, so a month after a published change is priced at the rate that applies to it. Every forecast is stored and scored against what actually happened, which is what turns next year's budget into something better than this year's guess. The attribution comes along too, so the annual line splits by product, team or customer, and an overrun has somewhere to be traced to rather than being announced.

Questions founders ask next

How do I budget for LLM costs for a year?

Build it month by month rather than multiplying a run rate, price each month at the rate in effect for it, and commit a range instead of a point. Then write down in advance what you would cut if the top of the range arrives, because that's the part that makes a budget useful rather than decorative.

How much can a published rate change move an annual budget?

In the modelled example, 33.3% on zero growth. A flat workload budgeted at $360,000 from a run rate actually costs $480,000 once Claude Sonnet 5's introductory pricing ends on 2026-09-01 and rises 50%, purely because eight of the twelve months fall after that date.

What is the difference between a forecast and a budget here?

A forecast is a guess you revise as evidence arrives. A budget is a commitment you're measured against. That makes a range more useful than a point for a budget, and it makes naming the trigger for action more useful than either.

Should the AI budget be one line?

Only if you never intend to cut it. A single company-wide line can't be defended in detail or reduced deliberately, because nobody can see which product or team it belongs to. Splitting it the way the business is run means an overrun has an owner and a first thing to try.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Keep reading


Sources: Anthropic pricing. Last reviewed 2026-08-05, rates effective 2026-07-02. Plain text version.