Guides / deepseek v4 flash pricing openrouter
DeepSeek V4 Flash pricing on OpenRouter
DeepSeek V4 Flash bills $0.14 per million input tokens and $0.28 output through OpenRouter, with cached input at $0.028. A routed rate carries a risk a direct one doesn't, because your invoice shows a router total rather than a line per model. Culpa, a local-first LLM cost, margin, and forecast ledger, prices each routed call from a dated price book so a change shows up.
Why this happens
When you call a provider directly, a price change eventually reaches you as an invoice you can read. Route through an aggregator and the invoice becomes one number covering every model you touched, so the per-model rate stops being something anybody checks. Culpa's own book proves how that goes wrong. It carried $0.089 and $0.180 per million for this model, and the live API reads $0.140 and $0.280. Nobody can now say whether the first pair was mis-read in July or the price rose since, because OpenRouter publishes no history. That's the whole problem in one row.
What this usually looks like
- Your router invoice is one line and your model mix is a dozen.
- Nobody can say what any single routed model costs you per million tokens.
- Router spend rose and the call count held steady, with no rate card to check against.
- The rate you budgeted with was copied once and never re-read.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Mistakes that cost the most
| Mistake | Why it hurts | Do instead |
|---|---|---|
| Budgeting a routed model from a rate you copied once. | A routed rate can move without an invoice line to announce it. | Re-read the router's own API on a schedule and date every rate you store. |
| Reading a router invoice as a per-model cost. | It's a total across every model you routed, so no single model's rate is visible in it. | Price each call from a dated book, then reconcile the sum against the router total. |
| Assuming a cached read costs a tenth here. | It's $0.028 against $0.14, a fifth rather than a tenth, so a cached line budgeted at a 90% discount costs twice what you planned. | Divide the cached rate by the input rate per model. That fraction is your real discount. |
Run this check tonight
- Read the router's models API and write down today's rate with today's date beside it.
- Compare that against whatever rate your budget or tooling currently holds.
- Price your own routed calls from the dated rate and sum them.
- Put that sum beside the router's invoice total for the same period.
- Repeat monthly. On a routed model the re-read is the only thing that catches a change.
What a rate that moved 57% costs before anybody notices
Illustrative example
A workload of 400M input and 120M output tokens a month on DeepSeek V4 Flash through OpenRouter, priced at the rate Culpa's book held until 2026-08-03 and at the rate the live API reports now. Volumes are modelled, both rate pairs are real.
A 57% rate move on unchanged traffic, invisible in a router total. The reconciliation is what turns it back into something you can see, and it needs a dated rate on your side of the wire.
Every number, with its confidence and source
| Figure | What it means | Confidence | Source |
|---|---|---|---|
| $0.14 per million | DeepSeek V4 Flash input rate through OpenRouter | calculated | openrouter.ai/api/v1/models reports 0.00000014 per input token, read 2026-08-03. Multiplied to a per-million figure. Seeded as migration 0034. |
| $57.20 to $89.60 | modelled monthly cost of one workload at the superseded rate and at the current one | estimated | Both endpoints from the teardown arithmetic at the two real dated rates. A range because the token volumes are modelled. |
What a generic answer can’t know
A router publishes today's rate and your invoice publishes one total. Neither tells you which of your calls went to which model, at what rate, on what day. That join lives only in your own traffic. Culpa keeps it on your infrastructure, holds your prompts and responses there, and counts the calls to run your plan.
Questions founders ask next
How much does DeepSeek V4 Flash cost on OpenRouter?
$0.14 per million input tokens and $0.28 output, with cached input at $0.028, read from OpenRouter's models API on 2026-08-03. Culpa's price book held $0.089 and $0.18 before that date and both rows are kept, so calls price at whichever rate was in force.
Why does Culpa keep the old rate as well?
Because nobody can prove it wrong. OpenRouter publishes no price history, so whether July was mis-transcribed or the price rose since is unknowable. Rewriting a rate Culpa can't vouch for would be inventing history, so the correction is dated instead.
Is the cached discount the same as on a frontier model?
No. Cached input here is $0.028 against $0.14 of uncached, a fifth of the input rate rather than a tenth. Anthropic and Gemini bill a tenth on every model. Groq bills about half where it offers a cached rate at all. The discount is set per model, and worth reading rather than assuming.
On your infrastructure
Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.
How Culpa works
Find the culprit. Not just the total.
Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.
Your prompts stay local.
Culpa runs on your own infrastructure. What you send to a model reaches us at no point.
Every dollar has a name.
Follow any charge to the conversation, the user, the feature and the customer behind it.
See the bill before it lands.
Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.
Three steps to your first answer.
Change one base URL.
Or drop in the Python or TypeScript library.
Find your most expensive conversation.
In the first session, not the first week.
Cost your next feature before you ship it.
Why the bill went up
Example dashboardCalls traced
418,209
across 3 projects
Spend this week
$378.41
+ $182 vs last week
Failed calls
312
74% retried, and you paid for all of them
+ $182 this week traced to one culprit
Spend over 14 days
Most expensive users
Next week forecast
Graded against reality. Accuracy shown as results land.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Keep reading
Sources: OpenRouter models. Last reviewed 2026-08-03, rates effective 2026-08-03. Plain text version.