Guides & answers

Gemini 3.8 Flash API Pricing Explained

·

Short answer

  • Gemini 3.8 Flash costs $0.75 / $3.75 per 1M input / output tokens — but that introductory rate expires December 31, 2026. On January 1, 2027 it becomes $1.50 / $7.50, and context caching doubles too.
  • The doubling applies to 3.6 and 3.7 Flash as well — all three share the same intro rate and the same expiry. The "standard" price is actually the old Gemini 3.5 Flash list price.
  • The per-task bill can exceed the sticker price: thinking tokens are billed as output tokens, and 3.8 Flash cannot drop below effort level low (no minimal setting). Google's own docs example burns 297 thought tokens to produce a 171-token answer.
  • Cheapest ways to soften it: Batch API at 50% off, effort level low, and context caching at $0.075/1M — each still doubling on Jan 1.

Every article about Gemini 3.8 Flash pricing says the same thing: $0.75 in, $3.75 out, doubles on January 1. That's correct and mostly useless, because it doesn't explain what the doubling actually is, or why two models at the exact same token price can produce very different invoices. Both answers are in Google's own docs, and both change what you should do before December 31.

What "doubles" actually means

Here's the part most explainers skip: the current price is a promotional rate, and the post-January price isn't a new number invented for 3.8 Flash. It's the old list price.

Compare the Flash lineup on Google's pricing page as of September 21, 2026 (paid tier, per 1M tokens, USD):

ModelInputOutputAfter Jan 1, 2027
Gemini 3.8 Flash$0.75$3.75$1.50 / $7.50
Gemini 3.7 Flash$0.75$3.75$1.50 / $7.50
Gemini 3.6 Flash$0.75$3.75$1.50 / $7.50
Gemini 3.5 Flash (earlier gen)$1.50$9.00no change — list price

Read the bottom row twice. 3.5 Flash — still listed, still sold at $1.50 / $9.00 with no expiry clause — never had an intro discount. What Google did across 3.6, 3.7, and 3.8 was run the whole new generation at roughly half the old list price and stamp every row with "through December 31, 2026." When the intro rate lapses on January 1, 2027, the Flash family simply returns to 3.5-era pricing — 3.8's doubled output rate ($7.50) actually lands below what 3.5 Flash charges today ($9.00).

So the January event is less "price hike" and more "coupon expires." That matters for planning: don't model 2027 costs as a shock; model 2026 as the discount window it always was.

The doubling is mechanical and model-wide: input, output, and context caching all carry the same "through Dec 31 / starting Jan 1" clauses, as does cached-token storage ($0.50 → $1.00 per 1M tokens per hour). Priority tier rises from $1.35 / $6.75 to $2.70 / $13.50. Only Batch and Flex keep their relative discount (50% of whatever standard becomes).

The thinking-token trap: same price, different bill

Now the part that actually breaks budgets. Google's pricing page bills Gemini 3.8 Flash output at "$3.75 per million output tokens (including thinking tokens)." The thinking docs are blunter: "response pricing is the sum of output tokens and thinking tokens." And one more sentence that deserves a highlighter: pricing is based on the full thought tokens the model generates, "despite only the summary being output from the API." You never see the reasoning; you pay for all of it.

Google's own streaming example in the thinking guide makes the ratio concrete. One modest logic puzzle (who lives in which colored house):

Token typeTokensBilled at
Input62input rate
Visible output171output rate
Thought tokens (invisible)297output rate

The reasoning consumed 74% more tokens than the answer itself — and on this class of task, thought tokens are the majority of what you pay for. 3.8 Flash is explicitly built to do this more, not less. Google's launch post says it plainly: the model "works harder," executing extra reasoning steps and iterative tool calls, and "might use more tokens to maximize performance, especially at higher effort levels." For efficiency-first workloads, Google's own suggestion is lower effort levels — or stay on 3.7 Flash.

One wrinkle makes 3.8 uniquely exposed: effort levels. All Flash models default to thinking medium, but 3.6 and 3.5 Flash can be dialed down to minimal. Gemini 3.8 Flash bottoms out at low — there is no minimal setting, so some thinking overhead is always billed. The floor is higher on the newest model.

A classic mistake: setting a small max_output_tokens to cap costs. The cap applies to thinking + output combined, so the model can exhaust it mid-reasoning, return a truncated or empty answer — and you still get billed for every thinking token generated. Google's documented advice is to lower thinking_level instead of shrinking max_output_tokens.

Four levers, ranked by effect

1. Batch mode — the only true 50% cut. Batch API jobs run at $0.375 / $1.875 per 1M tokens through December, $0.75 / $3.75 after. Anything that isn't user-facing (evals, backfills, document processing) belongs there. Trade-off: batch enqueued-token ceilings apply (3M for 3.8 Flash at the base tier) and results arrive when they arrive.

2. Effort level — the multiplier you control per request. Default is medium. For classification, extraction, and lookup-shaped work, low is what the docs recommend. It won't zero out thinking (that's impossible on 3.8), but it's the difference between a 300-token thought process and a 3,000-token one on complex prompts.

3. Context caching — cheap before it doubles, still cheap after. Cached input costs $0.075/1M now, $0.15 after January — a tenth of standard input either way, plus $0.50/1M/hr storage. If your agent resends the same system prompt and tool schemas every turn, this is the lever that matters. We've covered the same mechanics on DeepSeek's context caching, where cache hits cost $0.006/1M at peak — the concept transfers, the prices don't.

4. Model choice — the silent multiplier. 3.7 Flash sits at the identical token price with the identical January doubling, but it's the efficiency-tuned sibling ("high-speed, everyday coding" in Google's words) and supports lower effort settings. Same list price, leaner thoughts. And if the workload doesn't specifically need a Gemini, cheaper DeepSeek API options run Flash-class output at $1.20/1M peak — under a sixth of 3.8's doubled output rate — with off-peak discounts on top. At 3.8's intro price the gap is real; after January it becomes a canyon.

What to actually do before December 31

Three moves, in order of urgency:

Run the numbers now, not in December. Pull total_thought_tokens from your usage data (it's a first-class field on every response) and compute what January's rates do to your bill with your actual thought/output ratio. If thoughts are 60% of your output spend, your effective January output rate isn't $7.50 — it's closer to $12 per effective visible-answer token.

Shift everything non-interactive to Batch this quarter. The 50% Batch discount is percentage-based, so it survives the doubling — but every token you push through the standard tier before December 31 is paying 2026's premium for no reason.

Decide your January model mix explicitly. The free tier still exists in AI Studio (limited access, data used for training — fine for prototyping, not production; we've mapped the current free LLM API tiers here), and spend-based rate limits ($10 per rolling 10 minutes at Tier 1) will hard-stop runaway agent loops whether you planned for them or not.

All prices and dates above verified against Google's official pricing, thinking, and rate-limits documentation on September 21, 2026. Google adjusts pricing periodically — treat this as a snapshot of the intro-rate window, and re-check the pricing page before committing a 2027 budget.

Sources: Gemini Developer API pricing (Google, accessed Sep 21, 2026), Gemini thinking docs (Google), Gemini API rate limits (Google), Introducing Gemini 3.8 Flash (Google blog, Sep 2, 2026). DeepSeek comparison figures from DeepSeek's official models & pricing page (archived Sep 19, 2026).