Guides / culpa vs braintrust
Culpa vs Braintrust
Braintrust is an evaluation and observability platform billed on four separate meters at once. Culpa, a local-first LLM cost, margin, and forecast ledger, prices your provider calls from a versioned price book, sets them against what each customer pays, and forecasts next month. One bills you for tooling, the other measures what your product costs to run.
Why this happens
Braintrust's pricing page is unusually honest and unusually complicated, and both are worth saying. Starter is $0 a month with $10 of model credits, 1 GB of processed data then $4 per GB, 10k scores then $2.50 per 1k, and 14-day retention. Pro is $249 a month with $249 of credits, 5 GB then $3 per GB, 50k scores then $1.50 per 1k, and 30-day retention with archival at $0.50 per GB per month. That's a platform fee plus three independent meters, and the page even ships a calculator showing an estimated monthly total, which tells you the vendor knows the bill is hard to predict. Read that calculator carefully though, because it estimates what you'll pay BRAINTRUST. It isn't a forecast of what your models cost, and those are different numbers with different drivers. The 14-day retention on Starter is the other thing to note before building a cost practice on it.
What this usually looks like
- Your evals bill moved and you can't say which of the meters moved it.
- Processed data grew because traces got richer, not because usage grew.
- Your tooling bill and your model bill are both variable and neither predicts the other.
- Retention expired before anyone asked the question the traces would have answered.
- You can score a prompt version and not price it against what the customer pays.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Mistakes that cost the most
| Mistake | Why it hurts | Do instead |
|---|---|---|
| Reading a vendor's bill calculator as a spend forecast. | It estimates what you owe that vendor. Your provider bill has different drivers entirely. | Forecast the model spend separately, from your own call history. |
| Budgeting a multi-meter tool from its headline price. | The platform fee is one of four lines. Data volume and score count move independently of it. | Model each meter against your own traffic before committing to a tier. |
| Letting trace richness grow without watching processed data. | Adding fields to every span raises a per-GB meter that has nothing to do with model usage. | Treat trace payload size as a cost input, and measure it. |
Run this check tonight
- List every meter your tooling bills on, and which of them you actually watch.
- Check your retention window against the oldest question you regularly ask.
- Separate your tooling bill from your model bill and forecast them apart.
- Ask what a single customer costs you in model spend, against what they pay.
Two bills that move for unrelated reasons
Braintrust Pro publishes $249 a month including 5 GB of processed data, then $3 per GB. Take a team whose traces grow to 20 GB a month because they started logging retrieved documents, while their model spend is unchanged. Rates are published. The data volume is modelled.
Neither number is wrong and neither predicts the other. That's the argument for measuring model spend in a ledger that prices calls from a rate book, rather than inferring it from what your observability vendor charges you.
Every number, with its confidence and source
| Figure | What it means | Confidence | Source |
|---|---|---|---|
| $249 per month | Braintrust's published Pro platform fee, before its three usage meters | provider-reported | braintrust.dev/pricing, read 2026-08-03. Pro includes $249 of model credits, 5 GB processed data then $3 per GB, 50k scores then $1.50 per 1k, and 30-day retention plus archival at $0.50 per GB per month. Starter is $0 with $10 credits, 1 GB then $4 per GB, 10k scores then $2.50 per 1k, and 14-day retention. |
| $294 per month | the same Pro plan at a modelled 20 GB of processed data, with model spend unchanged | calculated | $249 platform fee plus 15 GB of overage above the 5 GB included, at the published $3 per GB: $249 + $45 = $294. The 20 GB monthly volume is modelled. The fee and the per-GB rate are published on braintrust.dev/pricing, read 2026-08-03. |
What a generic answer can’t know
Braintrust is a serious evaluation platform and this page doesn't pretend otherwise. It self-hosts on Enterprise, it publishes its meters plainly, and shipping a bill calculator is more transparency than most of this category offers. What it doesn't publish is margin against revenue or a forecast of your model spend, and its own numbers explain why: Braintrust meters what flows through Braintrust, so its view of your economics stops at its own invoice. Culpa reads your provider calls, prices each one from a versioned price book in exact decimal, carries the revenue each customer pays you so cost becomes margin, and forecasts next month with a range. Run the evaluator to know what ships. Run the ledger to know whether shipping it pays.
Questions founders ask next
How does Braintrust pricing work?
A platform fee plus three meters. Starter is $0 with $10 model credits, 1 GB processed data then $4 per GB, 10k scores then $2.50 per 1k, and 14-day retention. Pro is $249 with $249 credits, 5 GB then $3 per GB, 50k scores then $1.50 per 1k, and 30-day retention. Read 2026-08-03.
Does Braintrust forecast my LLM costs?
Its pricing page carries a calculator that estimates your monthly total to Braintrust, which is useful and is a different thing. Nothing in its published feature set claims to project your provider spend from your own history.
Do Braintrust and Culpa overlap?
Barely. Braintrust answers whether a version is good enough to ship, through evals and scores. Culpa answers what running it costs, what each customer pays against that, and what next month looks like. Neither replaces the other.
What's the retention on Braintrust's free tier?
14 days on Starter, 30 days on Pro with archival storage at $0.50 per GB per month, and custom on Enterprise, from its pricing page read 2026-08-03. Worth checking against the oldest cost question you ask, because a quarterly review needs more than either.
On your infrastructure
Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.
How Culpa works
Find the culprit. Not just the total.
Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.
Your prompts stay local.
Culpa runs on your own infrastructure. What you send to a model reaches us at no point.
Every dollar has a name.
Follow any charge to the conversation, the user, the feature and the customer behind it.
See the bill before it lands.
Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.
Three steps to your first answer.
Change one base URL.
Or drop in the Python or TypeScript library.
Find your most expensive conversation.
In the first session, not the first week.
Cost your next feature before you ship it.
Why the bill went up
Example dashboardCalls traced
418,209
across 3 projects
Spend this week
$378.41
+ $182 vs last week
Failed calls
312
74% retried, and you paid for all of them
+ $182 this week traced to one culprit
Spend over 14 days
Most expensive users
Next week forecast
Graded against reality. Accuracy shown as results land.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Keep reading
Sources: Braintrust, Braintrust pricing. Last reviewed 2026-08-03. Plain text version.