Guides / culpa vs athina

Culpa vs Athina

Athina is a collaborative AI development platform for building, testing and monitoring AI features, with 50-plus preset evaluations. Culpa, a local-first LLM cost, margin, and forecast ledger, prices each call from a versioned price book, attributes it to a customer against revenue, and forecasts next month with a range from your own history.

Why this happens

Start with what could and couldn't be verified, because on this page that's the finding. Athina publishes a detailed product: prompt management across models, 50-plus preset evaluations plus custom ones, dataset regeneration by swapping model, prompt or retriever, annotation for non-technical reviewers, and self-hosting. What it doesn't publish anywhere reachable is a price. The Pricing link in its own navigation returned a 404 on 2026-08-03, and a search across its homepage and documentation index, 33,400 characters in total, found no dollar figure at all. That's reported as an observation about what was reachable on the day rather than as a claim that no pricing exists, because the same conclusion was published wrongly about another vendor in this cluster after reading only part of a page. Either way, a tool whose cost you can't see is a tool you can't put in a forecast.

What this usually looks like

  • A tool is in your stack and its line on the budget is a guess.
  • You can evaluate a prompt thoroughly and not price the traffic it generates.
  • Procurement asks what a platform costs and the answer is a sales call.
  • Your model spend is forecast and your tooling spend isn't.
  • Nobody can say what a customer costs you across both.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run a free scan

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Assuming an unpublished price means an expensive one.It means unknown. Guessing high is as wrong as guessing low and both go into your plan as fact.Ask for the number in writing, and record it as unverified until you have it.
Leaving tooling out of the cost model because it's fixed.Most tools in this category meter something, so the fixed line is usually a floor.Model tooling and model spend together, and mark which parts are unverified.
Treating evaluation coverage as cost coverage.Fifty preset evals tell you about quality. None of them price a call against a rate book.Keep the evals and price the traffic separately.

Run this check tonight

  1. List every tool in your AI stack whose price you can state from memory.
  2. For the rest, mark the line as unverified rather than estimating it.
  3. Ask what your evaluation traffic itself costs in model spend.
  4. Check whether your forecast includes tooling or only provider bills.

What evaluation traffic costs when nobody prices it

Evaluation runs are model calls and they bill like model calls. A platform offering 50-plus preset evaluations invites you to run many of them, and LLM-as-judge evaluations call a model per scored output. Take a nightly regression over a modest dataset, priced on Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book, effective 2026-07-02. Volumes are modelled.

500 test cases x 6 evaluators = 3,000 judge calls per nightly run
3,000 x 30 nights = 90,000 judge calls a month
at 1,500 input and 100 output tokens each: 135M input and 9M output tokens
135 x $1.00 + 9 x $5.00 = $135.00 + $45.00 = $180.00 a month in judging alone

The evaluation harness has a provider bill of its own, and it grows with the number of evaluators rather than with user traffic. It rarely appears in anyone's cost model, which is exactly the kind of line a ledger that prices every call catches and a quality tool has no reason to.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
$180.00 per monthmodelled provider cost of a nightly LLM-as-judge regression, before any user trafficcalculated500 test cases x 6 evaluators x 30 nights = 90,000 judge calls, at 1,500 input and 100 output tokens each = 135M input and 9M output tokens. On Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book effective 2026-07-02: $135.00 + $45.00 = $180.00. Every volume here is modelled. The rates are real.

What a generic answer can’t know

Athina and Culpa aren't really competing for the same budget line, and the comparison is more useful read as a division of labour. Athina is where a team decides whether an AI feature is good enough: prompts, evaluations, datasets, annotation by domain experts who don't write code. Culpa is where the same team finds out what that feature costs to run, which customers it earns money from, and what next month looks like. Two specifics carry the difference. Culpa prices every call, including the calls your evaluation harness makes, from a versioned price book with effective dates. And it holds the revenue each customer pays you, so the answer is margin rather than spend. Neither appears in Athina's published feature set as read on 2026-08-03.

Questions founders ask next

What does Athina cost?

That couldn't be verified on 2026-08-03. The Pricing link in Athina's own navigation returned a 404, and no dollar figure appears across its homepage or documentation index, 33,400 characters searched in full. Recorded as unverified rather than as an absence of pricing.

Does Athina track LLM cost?

Nothing in its published material claims margin against revenue or a forecast of spend, as read on 2026-08-03. Its published strengths are prompt management, 50-plus preset and custom evaluations, dataset regeneration and annotation workflows for non-technical reviewers.

Does running evaluations cost money?

Yes, and it's routinely left out of cost models. LLM-as-judge evaluations are model calls billed at model rates, and they scale with the number of evaluators rather than with user traffic. A nightly regression across several evaluators can reach a meaningful monthly figure on its own.

Can I use Athina and Culpa together?

That's the natural shape. Athina decides whether the output is good enough. Culpa prices the traffic, including the evaluation traffic, sets it against what customers pay you, and forecasts the month.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run a free scan

Keep reading


Sources: Athina, Athina docs. Last reviewed 2026-08-03. Plain text version.