Guides / deepseek v4 flash pricing openrouter

DeepSeek V4 Flash pricing on OpenRouter

DeepSeek V4 Flash bills $0.14 per million input tokens and $0.28 output through OpenRouter, with cached input at $0.028. A routed rate carries a risk a direct one doesn't, because your invoice shows a router total rather than a line per model. Culpa, a local-first LLM cost, margin, and forecast ledger, prices each routed call from a dated price book so a change shows up.

Why this happens

When you call a provider directly, a price change eventually reaches you as an invoice you can read. Route through an aggregator and the invoice becomes one number covering every model you touched, so the per-model rate stops being something anybody checks. Culpa's own book proves how that goes wrong. It carried $0.089 and $0.180 per million for this model, and the live API reads $0.140 and $0.280. Nobody can now say whether the first pair was mis-read in July or the price rose since, because OpenRouter publishes no history. That's the whole problem in one row.

What this usually looks like

  • Your router invoice is one line and your model mix is a dozen.
  • Nobody can say what any single routed model costs you per million tokens.
  • Router spend rose and the call count held steady, with no rate card to check against.
  • The rate you budgeted with was copied once and never re-read.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run a free scan

Mistakes that cost the most

MistakeWhy it hurtsDo instead
Budgeting a routed model from a rate you copied once.A routed rate can move without an invoice line to announce it.Re-read the router's own API on a schedule and date every rate you store.
Reading a router invoice as a per-model cost.It's a total across every model you routed, so no single model's rate is visible in it.Price each call from a dated book, then reconcile the sum against the router total.
Assuming a cached read costs a tenth here.It's $0.028 against $0.14, a fifth rather than a tenth, so a cached line budgeted at a 90% discount costs twice what you planned.Divide the cached rate by the input rate per model. That fraction is your real discount.

Run this check tonight

  1. Read the router's models API and write down today's rate with today's date beside it.
  2. Compare that against whatever rate your budget or tooling currently holds.
  3. Price your own routed calls from the dated rate and sum them.
  4. Put that sum beside the router's invoice total for the same period.
  5. Repeat monthly. On a routed model the re-read is the only thing that catches a change.

What a rate that moved 57% costs before anybody notices

Illustrative example

A workload of 400M input and 120M output tokens a month on DeepSeek V4 Flash through OpenRouter, priced at the rate Culpa's book held until 2026-08-03 and at the rate the live API reports now. Volumes are modelled, both rate pairs are real.

At the old book rate: (400 x $0.089) + (120 x $0.18) = $35.60 + $21.60 = $57.20
At the rate read from the API 2026-08-03: (400 x $0.14) + (120 x $0.28) = $56.00 + $33.60 = $89.60
The same tokens cost $32.40 more a month, a 57% increase
Nothing in a router invoice separates that from a volume increase
Cached input at $0.028 against $0.14 is a 5x discount, not the 10x of a frontier model

A 57% rate move on unchanged traffic, invisible in a router total. The reconciliation is what turns it back into something you can see, and it needs a dated rate on your side of the wire.

Every number, with its confidence and source

FigureWhat it meansConfidenceSource
$0.14 per millionDeepSeek V4 Flash input rate through OpenRoutercalculatedopenrouter.ai/api/v1/models reports 0.00000014 per input token, read 2026-08-03. Multiplied to a per-million figure. Seeded as migration 0034.
$57.20 to $89.60modelled monthly cost of one workload at the superseded rate and at the current oneestimatedBoth endpoints from the teardown arithmetic at the two real dated rates. A range because the token volumes are modelled.

What a generic answer can’t know

A router publishes today's rate and your invoice publishes one total. Neither tells you which of your calls went to which model, at what rate, on what day. That join lives only in your own traffic. Culpa keeps it on your infrastructure, holds your prompts and responses there, and counts the calls to run your plan.

Questions founders ask next

How much does DeepSeek V4 Flash cost on OpenRouter?

$0.14 per million input tokens and $0.28 output, with cached input at $0.028, read from OpenRouter's models API on 2026-08-03. Culpa's price book held $0.089 and $0.18 before that date and both rows are kept, so calls price at whichever rate was in force.

Why does Culpa keep the old rate as well?

Because nobody can prove it wrong. OpenRouter publishes no price history, so whether July was mis-transcribed or the price rose since is unknowable. Rewriting a rate Culpa can't vouch for would be inventing history, so the correction is dated instead.

Is the cached discount the same as on a frontier model?

No. Cached input here is $0.028 against $0.14 of uncached, a fifth of the input rate rather than a tenth. Anthropic and Gemini bill a tenth on every model. Groq bills about half where it offers a cached rate at all. The discount is set per model, and worth reading rather than assuming.

On your infrastructure

Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.


How Culpa works

Find the culprit. Not just the total.

Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.

Your prompts stay local.

Culpa runs on your own infrastructure. What you send to a model reaches us at no point.

Every dollar has a name.

Follow any charge to the conversation, the user, the feature and the customer behind it.

See the bill before it lands.

Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.

Three steps to your first answer.

1

Change one base URL.

Or drop in the Python or TypeScript library.

2

Find your most expensive conversation.

In the first session, not the first week.

3

Cost your next feature before you ship it.

Base URLhttp://localhost:4545/v1Your traffic keeps flowing if Culpa ever stops.

Why the bill went up

Example dashboard

Calls traced

418,209

across 3 projects

Spend this week

$378.41

+ $182 vs last week

Failed calls

312

74% retried, and you paid for all of them

+ $182 this week traced to one culprit

Spend over 14 days

$0$20$40$60$8024262830020406
user_384report_generatorconv_91fprompt_v1894,220 tokens3 retries$6.81

Most expensive users

user_384$38.42
user_119$21.07
user_562$14.90
user_204$8.30
user_871$5.10

Next week forecast

Best$180
Median$240
p90$310
p99$395

Graded against reality. Accuracy shown as results land.

Free, no card, no account

Run the free Cost Leak Scan

It shows your most expensive conversation before you install anything.

Run a free scan

Keep reading


Sources: OpenRouter models. Last reviewed 2026-08-03, rates effective 2026-08-03. Plain text version.