# Culpa vs Athina > Athina's own pricing link returned a 404 on 2026-08-03 and no price appears across its site. Culpa publishes its rate book and forecasts your spend. URL: https://getculpa.com/culpa-vs-athina Last reviewed: 2026-08-03 ## Answer Athina is a collaborative AI development platform for building, testing and monitoring AI features, with 50-plus preset evaluations. Culpa, a local-first LLM cost, margin, and forecast ledger, prices each call from a versioned price book, attributes it to a customer against revenue, and forecasts next month with a range from your own history. ## Why this happens Start with what could and couldn't be verified, because on this page that's the finding. Athina publishes a detailed product: prompt management across models, 50-plus preset evaluations plus custom ones, dataset regeneration by swapping model, prompt or retriever, annotation for non-technical reviewers, and self-hosting. What it doesn't publish anywhere reachable is a price. The Pricing link in its own navigation returned a 404 on 2026-08-03, and a search across its homepage and documentation index, 33,400 characters in total, found no dollar figure at all. That's reported as an observation about what was reachable on the day rather than as a claim that no pricing exists, because the same conclusion was published wrongly about another vendor in this cluster after reading only part of a page. Either way, a tool whose cost you can't see is a tool you can't put in a forecast. ## What this usually looks like - A tool is in your stack and its line on the budget is a guess. - You can evaluate a prompt thoroughly and not price the traffic it generates. - Procurement asks what a platform costs and the answer is a sales call. - Your model spend is forecast and your tooling spend isn't. - Nobody can say what a customer costs you across both. ## Common mistakes - Assuming an unpublished price means an expensive one. Why it hurts: It means unknown. Guessing high is as wrong as guessing low and both go into your plan as fact. Do instead: Ask for the number in writing, and record it as unverified until you have it. - Leaving tooling out of the cost model because it's fixed. Why it hurts: Most tools in this category meter something, so the fixed line is usually a floor. Do instead: Model tooling and model spend together, and mark which parts are unverified. - Treating evaluation coverage as cost coverage. Why it hurts: Fifty preset evals tell you about quality. None of them price a call against a rate book. Do instead: Keep the evals and price the traffic separately. ## Self-check - List every tool in your AI stack whose price you can state from memory. - For the rest, mark the line as unverified rather than estimating it. - Ask what your evaluation traffic itself costs in model spend. - Check whether your forecast includes tooling or only provider bills. ## What evaluation traffic costs when nobody prices it Evaluation runs are model calls and they bill like model calls. A platform offering 50-plus preset evaluations invites you to run many of them, and LLM-as-judge evaluations call a model per scored output. Take a nightly regression over a modest dataset, priced on Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book, effective 2026-07-02. Volumes are modelled. 500 test cases x 6 evaluators = 3,000 judge calls per nightly run 3,000 x 30 nights = 90,000 judge calls a month at 1,500 input and 100 output tokens each: 135M input and 9M output tokens 135 x $1.00 + 9 x $5.00 = $135.00 + $45.00 = $180.00 a month in judging alone The evaluation harness has a provider bill of its own, and it grows with the number of evaluators rather than with user traffic. It rarely appears in anyone's cost model, which is exactly the kind of line a ledger that prices every call catches and a quality tool has no reason to. ## Cost figures Every figure carries its confidence and its source. No figure on this page is provider-reported. - $180.00 per month, modelled provider cost of a nightly LLM-as-judge regression, before any user traffic [calculated] Source: 500 test cases x 6 evaluators x 30 nights = 90,000 judge calls, at 1,500 input and 100 output tokens each = 135M input and 9M output tokens. On Claude Haiku 4.5 at real rates of $1.00 and $5.00 per million from the price book effective 2026-07-02: $135.00 + $45.00 = $180.00. Every volume here is modelled. The rates are real. ## FAQ Q: What does Athina cost? A: That couldn't be verified on 2026-08-03. The Pricing link in Athina's own navigation returned a 404, and no dollar figure appears across its homepage or documentation index, 33,400 characters searched in full. Recorded as unverified rather than as an absence of pricing. Q: Does Athina track LLM cost? A: Nothing in its published material claims margin against revenue or a forecast of spend, as read on 2026-08-03. Its published strengths are prompt management, 50-plus preset and custom evaluations, dataset regeneration and annotation workflows for non-technical reviewers. Q: Does running evaluations cost money? A: Yes, and it's routinely left out of cost models. LLM-as-judge evaluations are model calls billed at model rates, and they scale with the number of evaluators rather than with user traffic. A nightly regression across several evaluators can reach a meaningful monthly figure on its own. Q: Can I use Athina and Culpa together? A: That's the natural shape. Athina decides whether the output is good enough. Culpa prices the traffic, including the evaluation traffic, sets it against what customers pay you, and forecasts the month. ## Sources - Athina: https://www.athina.ai - Athina docs: https://docs.athina.ai Run the free Cost Leak Scan: https://app.getculpa.com/scan?source=pseo&slug=culpa-vs-athina&cluster=competitor Machine-readable index of every guide: https://getculpa.com/api/pages Human-readable index of every guide: https://getculpa.com/guides Site overview: https://app.getculpa.com/llms.txt Privacy: Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.