Guides / model switch cost calculator
Model switch cost calculator
This model switch cost calculator prices two models against the same token volumes and reports which one your own mix favours, plus the output-to-input ratio where the two tie. Culpa, a local-first LLM cost, margin, and forecast ledger, computes both totals in exact decimal, so a migration gets decided on arithmetic rather than on a headline rate.
Computed in your browser, in exact decimal
These two never cross. One is cheaper on both rates, so your token mix can't change the answer.
Rates come from Culpa's price book, effective 2026-07-02. Every figure is computed in exact decimal rather than floating point, which is why totals here match an invoice to the cent. Cost is calculated, not provider-reported: it prices the tokens you enter at a published rate.
Why this happens
A migration usually gets argued on one number, the input rate, because that's the column a shortlist has room for. Two models with different spreads between their input and output rates cross somewhere, and which side you sit on is decided by your own output-to-input ratio rather than by either rate card. Some pairs never cross at all, where one model is cheaper on both rates and the answer holds whatever your mix. Knowing which of those two situations you're in takes half a minute, and it decides whether the choice needs a meeting.
What this usually looks like
- A migration is being argued from input rates alone.
- Nobody has stated your output-to-input ratio for the workload in question.
- The saving was estimated once and never priced against real volumes.
- Two features with very different shapes are assumed to move together.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Mistakes that cost the most
| Mistake | Why it hurts | Do instead |
|---|---|---|
| Deciding a switch on the input rate. | The dearer output rate can reverse the answer well inside normal workloads. | Price both models at your real volumes. The totals settle it in one step. |
| Assuming every pair of models has a crossover. | Plenty of pairs never cross, and hunting for a switch point wastes the analysis. | Check whether one model is cheaper on both rates. If so, your mix changes nothing. |
| Migrating a whole product on one comparison. | Features sit at different ratios, so a single answer can be right and wrong at once. | Run the comparison per feature and let each land where its own arithmetic puts it. |
Run this check tonight
- Pull your uncached input, cached input and output volumes for one feature.
- Price your current model and the candidate against them above.
- Read the crossover the calculator reports and compare it against your own ratio.
- If your ratio sits near the crossover, the saving is small and the risk probably outweighs it.
- Repeat per feature rather than trusting one account-wide answer.
The same migration, two features, opposite answers
Illustrative example
Moving from GPT-5.4 nano at real rates of $0.20 per million input and $1.25 output to Llama 3.3 70B Versatile on Groq at $0.59 and $0.79, both effective 2026-07-02. Volumes are modelled.
The same migration costs $68.80 more on retrieval and saves $6.48 on drafting. One decision, two features, and the arithmetic disagrees with itself unless somebody runs it twice.
Every number, with its confidence and source
| Figure | What it means | Confidence | Source |
|---|---|---|---|
| 0.85 | output-to-input ratio at which the two teardown models cost the same | calculated | ($0.59 - $0.20) divided by ($1.25 - $0.79), using real GPT-5.4 nano and Groq Llama 3.3 70B Versatile rates per million from the price book, effective 2026-07-02. |
| $61.52 to $133.80 | modelled monthly cost across two features and two models, cheapest to dearest | estimated | The four totals in the teardown arithmetic at real rates. A range because the token volumes are modelled. |
What a generic answer can’t know
Both rate cards are public and your ratio per feature is yours alone, because it comes from token counts nobody else holds. A migration decided without it comes down to a headline. Culpa measures the volumes on your infrastructure, keeps your prompts and responses there, and counts the calls to run your plan.
Questions founders ask next
What does the crossover ratio mean?
It's the output-to-input ratio at which two models cost exactly the same. Below it the model with the cheaper input rate wins and above it the other one does, so the number tells you which side of the decision your own traffic sits on.
What if the calculator says the two never cross?
Then one model is cheaper on both input and output, and no token mix can change the ranking. That's the easy case, and it means the price question is settled before you start.
Should I switch on cost alone?
Cost is the part arithmetic settles in a minute, which is exactly why it's worth settling first. Quality on your own task needs a sample and someone's judgement, and it deserves the time that closing the price argument frees up.
On your infrastructure
Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.
How Culpa works
Find the culprit. Not just the total.
Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.
Your prompts stay local.
Culpa runs on your own infrastructure. What you send to a model reaches us at no point.
Every dollar has a name.
Follow any charge to the conversation, the user, the feature and the customer behind it.
See the bill before it lands.
Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.
Three steps to your first answer.
Change one base URL.
Or drop in the Python or TypeScript library.
Find your most expensive conversation.
In the first session, not the first week.
Cost your next feature before you ship it.
Why the bill went up
Example dashboardCalls traced
418,209
across 3 projects
Spend this week
$378.41
+ $182 vs last week
Failed calls
312
74% retried, and you paid for all of them
+ $182 this week traced to one culprit
Spend over 14 days
Most expensive users
Next week forecast
Graded against reality. Accuracy shown as results land.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Keep reading
Sources: OpenAI API pricing, Groq pricing. Last reviewed 2026-08-02. Plain text version.