Guides / culpa vs polarity

Culpa vs Polarity

Polarity post-trains a compact model inside your own infrastructure and routes routine work to it, keeping a frontier model for the hard cases. Culpa, a local-first LLM cost, margin, and forecast ledger, measures what that routing is worth by pricing every call before and after. One changes the bill and the other reads it, so they compete less than the shared vocabulary suggests.

Why this happens

Model optimisation and cost measurement get filed together because both promise a smaller bill, and they sit at opposite ends of the same job. Polarity's pitch is that most agent work repeats, so the repeating majority can run on a small model trained on your own traffic. That's a change to what you spend. Deciding which workflow to move needs a measurement first, because the candidate is whichever repeating path costs most, and confirming the move worked needs the same measurement afterwards on the same traffic. Without that, a migration is judged on a vendor's number rather than on your own.

What this usually looks like

  • A model migration is being considered and nobody can rank workflows by what they actually cost.
  • Nobody can name which repeating path carries the most spend.
  • A previous optimisation was declared a success without a before-and-after on your own traffic.
  • Frontier-model spend is assumed to be necessary because nobody has measured the routine share.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run a free scan

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Choosing what to optimise from intuition.The workflow that feels expensive and the one that really is are often different.Rank workflows by measured spend first. The candidate list usually reorders.
Accepting a published saving as your saving.A vendor's customer result was measured on that customer's traffic and mix, not yours.Price your own before-and-after on the same workload. Treat any published figure as a hypothesis.
Treating optimisation as a one-off.Traffic mix drifts, so a routing decision that paid last quarter can stop paying quietly.Keep measuring after the migration. The saving is a rate, not an event.

Run this check tonight

  1. List your workflows and rank them by measured monthly spend rather than by feel.
  2. For the top one, work out what share of its calls are routine and repeating.
  3. That share is the part any routing approach can address, and it's the size of the prize.
  4. Price the workload before the change, then price the identical workload after.
  5. Re-check quarterly, because the mix that justified the move keeps moving.

Why the measurement has to come first and last

Polarity's published approach, read from polarity.so on 2026-08-03, against what Culpa contributes to the same decision. No pricing comparison is drawn, because Polarity's pricing is demo-led and not published, so any figure would be invented.

Polarity: identifies the repeating majority of a workflow and routes it to a compact model
Polarity: post-trains that model inside the customer's own infrastructure and keeps learning from production
Polarity publishes a stat block reading 72% lower inference cost, 2.4x faster responses and +8% accuracy over Opus 4.7
A separate testimonial says the bill was cut by nearly three quarters and accuracy held. The site names no company
So the figures are the vendor's published claims rather than a measurement on your traffic or Culpa's
Culpa before: ranks workflows by measured spend, so the migration candidate is chosen rather than guessed
Culpa after: prices the identical workload again, so the saving is your number rather than a published one

Nothing here is a head-to-head. A routing change without a before-and-after on your own traffic is a decision taken on somebody else's evidence, and that's the gap a ledger fills at both ends.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
72% lower inference costa customer result Polarity publishes on its own site, not a Culpa measurementprovider-reportedpolarity.so, read 2026-08-03, where an unattributed stat block reads 72% lower inference cost, 2.4x faster responses and +8% accuracy over Opus 4.7. A separate testimonial from Anton Reza, CTO, says the bill was cut by nearly three quarters and accuracy held. The site names no company, so whose traffic produced the figures goes unstated.

What a generic answer can’t know

What share of your traffic is genuinely routine is a fact about your product that no vendor can know in advance, and it decides whether any routing approach is worth the work. It comes out of your own calls. Culpa measures them on your infrastructure, keeps your prompts and responses there, and counts the calls to run your plan.

Questions founders ask next

Is Polarity a Culpa competitor?

Less than the category suggests. Polarity changes which model runs your routine work and Culpa measures what everything costs. The honest framing is that a routing decision needs a measurement on both sides of it, which is the part Culpa does.

Does Culpa reduce my bill?

Not by itself. Culpa finds where the money goes and forecasts where it's heading, and the reductions come from what you do with that, whether that's a routing change, a cache, a prompt or a model swap. Culpa's job is making the number you act on a real one.

Is Polarity's published 72% saving something I'd get?

Treat it as a hypothesis rather than a forecast. It's a stat Polarity publishes on its own site, sitting near a testimonial that says a bill was cut by nearly three quarters. The site names no company and no methodology, so whose traffic produced it isn't stated. Your own share of routine work decides your number, and that's measurable before you commit.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run a free scan

Keep reading


Sources: Polarity. Last reviewed 2026-08-03. Plain text version.