Guides / culpa vs polarity
Culpa vs Polarity
Polarity post-trains a compact model inside your own infrastructure and routes routine work to it, keeping a frontier model for the hard cases. Culpa, a local-first LLM cost, margin, and forecast ledger, measures what that routing is worth by pricing every call before and after. One changes the bill and the other reads it, so they compete less than the shared vocabulary suggests.
Why this happens
Model optimisation and cost measurement get filed together because both promise a smaller bill, and they sit at opposite ends of the same job. Polarity's pitch is that most agent work repeats, so the repeating majority can run on a small model trained on your own traffic. That's a change to what you spend. Deciding which workflow to move needs a measurement first, because the candidate is whichever repeating path costs most, and confirming the move worked needs the same measurement afterwards on the same traffic. Without that, a migration is judged on a vendor's number rather than on your own.
What this usually looks like
- A model migration is being considered and nobody can rank workflows by what they actually cost.
- Nobody can name which repeating path carries the most spend.
- A previous optimisation was declared a success without a before-and-after on your own traffic.
- Frontier-model spend is assumed to be necessary because nobody has measured the routine share.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Mistakes that cost the most
| Mistake | Why it hurts | Do instead |
|---|---|---|
| Choosing what to optimise from intuition. | The workflow that feels expensive and the one that really is are often different. | Rank workflows by measured spend first. The candidate list usually reorders. |
| Accepting a published saving as your saving. | A vendor's customer result was measured on that customer's traffic and mix, not yours. | Price your own before-and-after on the same workload. Treat any published figure as a hypothesis. |
| Treating optimisation as a one-off. | Traffic mix drifts, so a routing decision that paid last quarter can stop paying quietly. | Keep measuring after the migration. The saving is a rate, not an event. |
Run this check tonight
- List your workflows and rank them by measured monthly spend rather than by feel.
- For the top one, work out what share of its calls are routine and repeating.
- That share is the part any routing approach can address, and it's the size of the prize.
- Price the workload before the change, then price the identical workload after.
- Re-check quarterly, because the mix that justified the move keeps moving.
Why the measurement has to come first and last
Polarity's published approach, read from polarity.so on 2026-08-03, against what Culpa contributes to the same decision. No pricing comparison is drawn, because Polarity's pricing is demo-led and not published, so any figure would be invented.
Nothing here is a head-to-head. A routing change without a before-and-after on your own traffic is a decision taken on somebody else's evidence, and that's the gap a ledger fills at both ends.
Every number, with its confidence and source
| Figure | What it means | Confidence | Source |
|---|---|---|---|
| 72% lower inference cost | a customer result Polarity publishes on its own site, not a Culpa measurement | provider-reported | polarity.so, read 2026-08-03, where an unattributed stat block reads 72% lower inference cost, 2.4x faster responses and +8% accuracy over Opus 4.7. A separate testimonial from Anton Reza, CTO, says the bill was cut by nearly three quarters and accuracy held. The site names no company, so whose traffic produced the figures goes unstated. |
What a generic answer can’t know
What share of your traffic is genuinely routine is a fact about your product that no vendor can know in advance, and it decides whether any routing approach is worth the work. It comes out of your own calls. Culpa measures them on your infrastructure, keeps your prompts and responses there, and counts the calls to run your plan.
Questions founders ask next
Is Polarity a Culpa competitor?
Less than the category suggests. Polarity changes which model runs your routine work and Culpa measures what everything costs. The honest framing is that a routing decision needs a measurement on both sides of it, which is the part Culpa does.
Does Culpa reduce my bill?
Not by itself. Culpa finds where the money goes and forecasts where it's heading, and the reductions come from what you do with that, whether that's a routing change, a cache, a prompt or a model swap. Culpa's job is making the number you act on a real one.
Is Polarity's published 72% saving something I'd get?
Treat it as a hypothesis rather than a forecast. It's a stat Polarity publishes on its own site, sitting near a testimonial that says a bill was cut by nearly three quarters. The site names no company and no methodology, so whose traffic produced it isn't stated. Your own share of routine work decides your number, and that's measurable before you commit.
On your infrastructure
Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.
How Culpa works
Find the culprit. Not just the total.
Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.
Your prompts stay local.
Culpa runs on your own infrastructure. What you send to a model reaches us at no point.
Every dollar has a name.
Follow any charge to the conversation, the user, the feature and the customer behind it.
See the bill before it lands.
Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.
Three steps to your first answer.
Change one base URL.
Or drop in the Python or TypeScript library.
Find your most expensive conversation.
In the first session, not the first week.
Cost your next feature before you ship it.
Why the bill went up
Example dashboardCalls traced
418,209
across 3 projects
Spend this week
$378.41
+ $182 vs last week
Failed calls
312
74% retried, and you paid for all of them
+ $182 this week traced to one culprit
Spend over 14 days
Most expensive users
Next week forecast
Graded against reality. Accuracy shown as results land.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Keep reading
Sources: Polarity. Last reviewed 2026-08-03. Plain text version.