Guides / gemini 3.5 flash pricing
Gemini 3.5 Flash pricing, and why the version bump costs
Gemini 3.5 Flash pricing is $1.50 per million input tokens and $9.00 per million output tokens, five times the input rate of Gemini 2.5 Flash at $0.30. A version bump inside the same family multiplies your bill rather than trimming it. Culpa, a local-first LLM cost, margin, and forecast ledger, prices the same workload on both, so the move stays a decision.
Why this happens
Version numbers carry an assumption that newer means better and cheaper, because that's how software usually goes. Model families break it. Within Gemini's own Flash line the 3.5 release prices input at five times the 2.5 release and output at 3.6 times. Nothing on the rate card flags that as an increase, because a rate card shows levels rather than changes, and a team upgrading on the strength of a release note finds out at the end of the month.
What this usually looks like
- An upgrade shipped on quality evidence with no cost comparison beside it.
- Spend rose sharply in the month a model version changed and the two were treated as separate events.
- Nobody priced the outgoing model and the incoming one against the same workload.
- The release note was read and the rate card wasn't.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Mistakes that cost the most
| Mistake | Why it hurts | Do instead |
|---|---|---|
| Assuming a newer version of the same family costs the same or less. | Here it costs five times more on input, and nothing in the naming warns you. | Price both versions against last month's real token volumes before scheduling the migration. |
| Migrating everything at once. | The whole workload moves to the higher rate to fix the share that needed the newer model. | Move the traffic that benefits and leave the rest, then measure the quality difference on each. |
| Judging the upgrade on quality alone. | At 4.15 times the cost the newer model has to be a great deal better, not slightly better. | Set the bar at the cost multiple. If it can't clear that, the upgrade is a preference. |
Run this check tonight
- Take last month's token volumes and price them on 2.5 Flash and 3.5 Flash side by side.
- Divide the two totals. That multiple is the bar the quality improvement has to clear.
- Check what share of your traffic actually benefits from the newer model.
- Price a split where only that share moves and the rest stays.
- Compare both against Gemini 3.1 Flash-Lite, which undercuts even the older Flash on input.
The same month, on two versions of Flash
Illustrative example
A workload of 80 million input and 15 million output tokens in a month, priced on Gemini 2.5 Flash at $0.30 and $2.50 per million and on 3.5 Flash at real rates. Volumes are modelled.
A full migration costs 4.15 times the old bill. Moving only the fifth of traffic that benefits costs 1.6 times, and that difference is the whole return on doing the analysis.
Every number, with its confidence and source
| Figure | What it means | Confidence | Source |
|---|---|---|---|
| $1.50 / $9.00 per million | Gemini 3.5 Flash input and output rates | calculated | Culpa price book row for gemini/gemini-3.5-flash, $0.0015 and $0.009 per 1k, effective 2026-07-02, verified against Gemini API pricing. Shown here per million. |
| 5x | Gemini 3.5 Flash input rate as a multiple of Gemini 2.5 Flash | calculated | $1.50 divided by $0.30, from the gemini-3.5-flash and gemini-2.5-flash price-book rows. |
| $61.50 to $255.00 | modelled monthly cost of one identical workload on 2.5 Flash against 3.5 Flash | estimated | Both endpoints from the teardown arithmetic at real rates for each model. A range because the token volumes are modelled. |
What a generic answer can’t know
A rate card shows both models and never shows the change between them, because it has no memory of what you were running last month. Which share of your traffic genuinely benefits is a fact about your own prompts. Culpa holds both on your infrastructure, keeps your prompts and responses there, and counts the calls to run your plan.
Questions founders ask next
How much does Gemini 3.5 Flash cost?
$1.50 per million input tokens and $9.00 per million output, with context caching at $0.15 per million. Rates verified against Google's pricing page.
Is Gemini 3.5 Flash more expensive than 2.5 Flash?
Substantially. Input costs five times more, $1.50 against $0.30, and output 3.6 times more, $9.00 against $2.50. On a typical token mix the whole bill lands around 4.15 times higher.
Should I upgrade from 2.5 Flash?
Only for the traffic that measurably benefits. At four times the cost the quality gain has to be large rather than marginal, and a partial migration of the fifth of traffic that needs it usually beats moving everything.
On your infrastructure
Culpa runs on your infrastructure. Your prompts and responses never leave it. Culpa counts calls to run your plan, and it fails open, so if it ever breaks your app keeps running.
How Culpa works
Find the culprit. Not just the total.
Your dashboard shows what you spent. It stops short of who spent it. Culpa shows the conversation, the user and the feature behind it.
Your prompts stay local.
Culpa runs on your own infrastructure. What you send to a model reaches us at no point.
Every dollar has a name.
Follow any charge to the conversation, the user, the feature and the customer behind it.
See the bill before it lands.
Cost your next feature before you ship it. You get the likely bill and the worst case, at best, median, p90 and p99.
Three steps to your first answer.
Change one base URL.
Or drop in the Python or TypeScript library.
Find your most expensive conversation.
In the first session, not the first week.
Cost your next feature before you ship it.
Why the bill went up
Example dashboardCalls traced
418,209
across 3 projects
Spend this week
$378.41
+ $182 vs last week
Failed calls
312
74% retried, and you paid for all of them
+ $182 this week traced to one culprit
Spend over 14 days
Most expensive users
Next week forecast
Graded against reality. Accuracy shown as results land.
Free, no card, no account
Run the free Cost Leak Scan
It shows your most expensive conversation before you install anything.
Keep reading
Sources: Gemini API pricing. Last reviewed 2026-08-01, rates effective 2026-07-02. Plain text version.