Google's current model docs recommend both Gemini 3.8 Flash and Gemini 3.5 Flash-Lite for new projects. That's not indecision on Google's part: it's two models built for different jobs that happen to sit next to each other in the same API. If you're starting a new integration and just want to know which model code to type into client.interactions.create(), "both" isn't an answer you can ship with.
I've spent the last couple weeks running the same workloads through both models (extraction jobs, a multimodal video-tagging pipeline, some agentic coding tasks) to figure out where the line actually is. This isn't a rehash of Gemini prompting fundamentals (we've got a full Gemini 3.8 Flash guide for that, covering video prompting, grounding, and code execution in depth). This is specifically the "which one do I call" decision, with the numbers that should drive it.
The spec sheet, side by side
| Gemini 3.8 Flash | Gemini 3.5 Flash-Lite | |
|---|---|---|
| Model code | gemini-3.8-flash | gemini-3.5-flash-lite |
| Input limit | 1,048,576 tokens | 1,048,576 tokens |
| Output limit | 65,536 tokens | 65,536 tokens |
| Modalities | text, image, video, audio, PDF | text, image, video, audio, PDF |
| Thinking levels | low, medium (default), high | minimal (default), low, medium, high |
| Input price | $0.75/1M through Dec 31, 2026; $1.50/1M from Jan 1, 2027 | $0.30/1M, stable |
| Output price | $3.75/1M through Dec 31, 2026; $7.50/1M from Jan 1, 2027 | $2.50/1M, stable |
What jumps out immediately: the context window and modality support are identical. This isn't a "big model vs small model" split in the way GPT-4o vs GPT-4o-mini used to be. It's a reasoning-depth split. 3.8 Flash can't go below low thinking: Google explicitly blocks minimal and throws an error if you try it. 3.5 Flash-Lite defaults to minimal and only goes up from there if you ask.
That one detail tells you almost everything about how Google wants you to use these two models. 3.8 Flash is built to think, even at its cheapest setting. 3.5 Flash-Lite is built to not think unless you tell it to.
The actual decision framework
I ran four categories of task through both models: structured extraction, multimodal classification, agentic coding, and bulk/high-volume processing. Here's where each one won.
Extraction (invoices, forms, structured JSON from text)
This is the closest call, and it depends entirely on how gnarly your source documents are.
For clean, consistent extraction, pulling {name, date, amount} out of a well-formatted invoice template you've seen a thousand times, 3.5 Flash-Lite at minimal thinking is the right default. You're paying $0.30/1M input against $0.75/1M (or worse, $1.50/1M after the promo ends), and there's no reasoning task here worth burning tokens on. The model doesn't need to "think" about where the total is on a template it's seen before.
Where I switched to 3.8 Flash: multi-page documents with inconsistent layouts, nested tables, or fields that require cross-referencing (e.g., "does the shipping address match the billing address, and if not, flag it"). That's a small reasoning step, and 3.5 Flash-Lite at minimal will hallucinate a match more often than 3.8 Flash at low. If you're extracting from messy real-world PDFs (scanned contracts, vendor invoices that don't follow a template) bump to 3.8 Flash and set thinking_level: "low". The cost delta at low thinking is small enough that it's not worth the accuracy hit.
Rule of thumb: if your extraction schema has a field that requires comparing two other fields, use 3.8 Flash. If every field is a direct lookup, use 3.5 Flash-Lite.
Multimodal (image/video classification, tagging, description)
Both models accept the same modalities, which surprised me: I expected Flash-Lite to be text-only or image-only. It isn't. It'll take video and audio input just like 3.8 Flash.
But multimodal reasoning is where minimal thinking shows its limits fastest. I ran a video-tagging job (classify the primary activity in 10-second clips into one of 40 categories) through both. 3.5 Flash-Lite at minimal was fast and cheap but confused visually similar categories (cycling vs. spin class, cooking vs. baking) at a noticeably higher rate than 3.8 Flash at low. Bumping Flash-Lite to medium thinking closed most of that gap, but at that point you're paying reasoning-model latency for a "lite" model and the price gap to 3.8 Flash narrows enough that it's not obviously the better trade.
For straightforward multimodal tasks, "is there a person in this image," "transcribe this audio," "what's the dominant color": Flash-Lite at minimal is fine and notably faster. For multimodal tasks requiring any disambiguation, default to 3.8 Flash.
Agentic coding and multi-step tool use
Not close. 3.8 Flash is explicitly the one Google built for "long-horizon software engineering, autonomous agents, and complex enterprise workflows," and it shows. Multi-step tool-calling loops, read a file, decide what to change, write it, run tests, interpret the failure, retry: need the model to hold state across turns and reason about what went wrong. 3.5 Flash-Lite can technically do this at high thinking, but at that point you've erased its cost advantage and you're just running a worse version of 3.8 Flash.
If you're building anything that resembles an agent loop, see our reasoning models guide for how thinking budgets interact with agentic tasks generally: start with 3.8 Flash at medium (its default) and only drop to low if latency is the bottleneck.
High-volume, low-complexity batch jobs
This is Flash-Lite's home turf, and it's not close either. Sentiment tagging on support tickets, spam/not-spam classification, simple content moderation flags, summarizing short text, anything where you're running hundreds of thousands or millions of calls and each individual call is a shallow judgment.
At minimal thinking, Flash-Lite's $0.30/$2.50 pricing against 3.8 Flash's $0.75/$3.75 (current promo) means Flash-Lite runs roughly 40-60% cheaper per call depending on your input/output ratio, with output tokens weighted more heavily since that's where the price gap is widest. On a million-call batch job, that's a real budget line, not a rounding error. And once 3.8 Flash's standard pricing kicks in on January 1, 2027 ($1.50/$7.50), the gap gets wider, not narrower.
If you're not sure whether your batch job is "simple enough" for Flash-Lite, that's what evals are for: run a sample through both and compare accuracy before committing a million-call budget to either one. Our guide to testing prompts with promptfoo covers the harness for exactly this kind of before-you-commit comparison.
A quick decision checklist
Ask these in order:
- Does any part of the task require comparing, cross-referencing, or inferring beyond a direct lookup? If yes, use 3.8 Flash.
- Is this an agent loop with multiple tool calls per task? If yes, use 3.8 Flash: don't bother testing Flash-Lite first.
- Is this a single-shot classification or extraction on a large volume of similar inputs? If yes, start with 3.5 Flash-Lite at
minimaland only raise the thinking level if your eval accuracy is below target. - Still unsure? Default to 3.5 Flash-Lite at
minimal, run your eval set, and upgrade tolow/mediumthinking or to 3.8 Flash only where the numbers say you need it. It's cheaper to start low and prove you need more than to start high and never check if you could've spent less.
The SDK detail that trips people up
If you're pulling code samples from older tutorials, watch for this: the current google-genai SDK pattern is
from google import genai
client = genai.Client()
response = client.interactions.create(
model="gemini-3.8-flash",
input="Summarize the key findings in two sentences.",
generation_config={"thinking_level": "low"},
)
print(response.output_text)
This replaced the older google.generativeai package's GenerativeModel(...).generate_content(...) pattern. A lot of blog posts and Stack Overflow answers still use the deprecated import: if your code throws an AttributeError on generate_content, that's almost always why. Same applies to thinking_level: it's a generation_config key now, not a separate parameter or a model-name suffix.
Where to go deeper
This post is deliberately narrow: it's the "which model" decision, not a Gemini prompting tutorial. For thinking-level tuning, video prompting patterns, grounding with search, and code execution specifics, the Gemini prompting guide covers 3.8 Flash in full. If you're weighing Gemini against other providers entirely, the model comparison guide has the wider picture.
One caveat worth stating plainly: pricing and default thinking levels are the kind of thing Google changes on short notice, 3.8 Flash was Google's third Flash release in six weeks as of this writing. Check ai.google.dev/gemini-api/docs/pricing before you lock in a production budget; treat the numbers here as accurate as of September 23, 2026, not as a permanent contract.



