Guides / ai cost in investor diligence

The AI cost questions an investor diligence process actually asks

Diligence tests whether your unit economics survive growth, which is a different question from whether this month looked fine. The answers need cost joined to customers, and concentration is the one most founders can't produce. Culpa, a local-first LLM cost, margin, and forecast ledger, holds cost and revenue together so those answers exist before they're asked.

Why this happens

A board wants to know how the quarter went. An investor wants to know what happens at ten times the volume, and those questions have different answers from the same data. The one that catches people out is concentration. Model spend is rarely spread evenly across customers, and a small group usually accounts for a large share of it. That's fine, and it's fine in a way that needs demonstrating rather than asserting, because the follow-up question is whether your growth is coming from the expensive group. If it does, blended margin falls as you scale, which is the opposite of what a software business is supposed to do and exactly the thing being tested. The reason this goes badly isn't that the numbers are bad. It's that the join between model spend and customer identity often doesn't exist, so the honest answer is that nobody knows, and in a diligence room that reads as worse than a difficult number.

What this usually looks like

  • You can state total AI spend and not spend by customer.
  • Nobody knows what share of model cost your top decile accounts for.
  • Gross margin is reported blended and never by segment.
  • Your fastest-growing cohort's unit cost has never been calculated.
  • The AI line appears in the model as a percentage of revenue with no derivation.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run the free Cost Leak ScanStart 14-day trial

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Bringing total spend to a diligence conversation.The question is whether economics hold at scale, and a total says nothing about that.Bring cost per customer, its distribution, and margin by segment.
Reporting blended gross margin only.It hides whether the growing segment is the profitable one, which is the actual test.Split margin by cohort and by plan, and show the trend in each.
Modelling AI cost as a fixed percentage of revenue.It assumes the thing being questioned, and a reader will notice the assumption immediately.Derive it from token volumes and rates, and show the derivation.
Answering concentration questions from memory.The gap between an estimate and a query is obvious, and it costs credibility on everything else.Produce it from the ledger, with the method stated.

Run this check tonight

  1. Rank your customers by model spend and work out what share the top decile accounts for.
  2. Calculate gross margin for that decile and for everyone else, separately.
  3. Check which group your last two quarters of growth came from.
  4. Try to answer all three from a query rather than a spreadsheet, and see whether you can.

What cost concentration looks like when you measure it

Illustrative example

A modelled book of 1,000 customers and $50,000 of monthly model spend, where the top 100 customers account for 60% of it. The concentration, the customer count and the spend are all modelled, and the point is the ratio rather than the amounts.

top 100 customers: 60% of $50,000 = $30,000, which is $300.00 each
remaining 900 customers: $20,000, which is $22.22 each
the heavy decile costs 13.5 times the rest, per customer
if the next 1,000 customers resemble the heavy group, spend rises $300,000 rather than $22,222

Thirteen and a half times isn't a problem by itself. Not knowing the number is. The follow-up question in every diligence conversation is which group your growth is coming from, and that only has an answer if cost was attributed to customers while it was being incurred.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
13.5xmodelled cost per customer in the top decile against everyone elsecalculatedA modelled 1,000 customers and $50,000 of monthly model spend with 60% concentrated in the top 100, which is $30,000 across 100 customers, or $300.00 each, against $20,000 across 900, or $22.22 each, a ratio of 13.5. Every figure is modelled, and the ratio rather than the amounts is the point.

What a generic answer can’t know

The join between model spend and a customer is the whole difficulty, and it has to be made at the call. A provider invoice records tokens against keys and models, your billing system records revenue against customers, and nothing connects them by default, which is why cost per customer is so often estimated and so rarely queried. Culpa prices every call in exact decimal from a versioned price book, attributes it to the customer, plan and feature that caused it, and takes your revenue alongside, so concentration and margin by segment are queries with a method you can state. Forecasts are stored as ranges and scored against what happened, which is what lets a projection be defended rather than presented. In a room testing whether the economics hold, being able to run the query is most of the answer.

Questions founders ask next

What do investors ask about AI costs?

Whether unit economics hold as volume grows. In practice that becomes three questions: cost per customer and its distribution, gross margin by segment, and which segment your growth is coming from. All three need model spend joined to customer identity.

Why does cost concentration matter?

Because it decides what happens to margin as you scale. In the modelled example the top decile costs $300.00 a customer against $22.22 for everyone else, 13.5 times. If growth comes from the expensive group, blended margin falls as the business grows.

Is a high concentration bad?

Not on its own, and plenty of healthy businesses have one. What matters is knowing the number and whether those customers are priced for it. The damaging answer in a diligence conversation isn't a difficult figure, it's having none.

What if we can't produce cost per customer?

Then that's the finding, and it's worth surfacing before someone else does. The join has to be made when the call happens, so it can't be reconstructed for past periods. Starting now means the next quarter has it even if the last one doesn't.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Keep reading


Sources: Anthropic pricing. Last reviewed 2026-08-05, rates effective 2026-07-02. Plain text version.