Gemini's standout feature is what it can take in. Gemini 3.8 Flash accepts text, images, video, audio and PDFs in a single request, with a 1,048,576-token input window. It's the only one of the big three frontier model families that takes video directly, and that changes what you can build: meeting analysis, video QA, screen-recording bug reports.
This guide covers how to prompt Gemini 3.8 Flash and 3.5 Flash-Lite through the current google-genai SDK.
The current Gemini lineup
Google's recommendation for new projects is simple: use 3.8 Flash or 3.5 Flash-Lite.
| Gemini 3.8 Flash | Gemini 3.5 Flash-Lite | |
|---|---|---|
| Model code | gemini-3.8-flash | gemini-3.5-flash-lite |
| Input limit | 1,048,576 tokens | 1,048,576 tokens |
| Output limit | 65,536 tokens | 65,536 tokens |
| Inputs | Text, image, video, audio, PDF | Text, image, video, audio, PDF |
| Thinking levels | low, medium, high | minimal (default), low, medium, high |
| Price (per 1M tokens, in / out) | $0.75 / $3.75 until Dec 31, 2026, then $1.50 / $7.50 | $0.30 / $2.50 |
Use 3.8 Flash as your default. Drop to Flash-Lite for classification, extraction and other high-volume routes once your tests show it holds up. Other current models are compared on the models page.
If you're migrating code: the old google.generativeai package and GenerativeModel(...).generate_content(...) pattern are what most 2024–2025 tutorials show. Current docs use google-genai with a client:
from google import genai
client = genai.Client() # reads GEMINI_API_KEY from the environment
interaction = client.interactions.create(
model="gemini-3.8-flash",
system_instruction="You are a senior data analyst. Be concise and specific.",
input="Explain the difference between median and mean income in two sentences.",
)
print(interaction.output_text)
Controlling thinking
Gemini 3.x models reason before answering. You set how much with thinking_level:
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Plan a zero-downtime migration from MySQL 5.7 to Postgres 17 for a 2TB database.",
generation_config={"thinking_level": "high"},
)
Google's own guidance on levels:
minimal/low: simple fact retrieval (minimalisn't available on 3.8 Flash)medium: moderate reasoning; the default on most Gemini 3.x modelshigh: complex coding and math
Flash-Lite defaults to minimal, which suits its high-volume role. If a Flash-Lite route starts failing on anything that needs reasoning, try low or medium before switching to a bigger model.
As with other reasoning models, don't add "think step by step." Describe the task, the constraints and the output format, and let the thinking level set depth. The raw thoughts aren't returned; thought summaries are available if you need to audit the logic. See how to prompt reasoning models.
Prompting video
This is where Gemini is in a category of its own. Upload the file, wait for processing, then ask:
import time
from google import genai
client = genai.Client()
video = client.files.upload(file="product-meeting.mp4")
while not video.state or video.state.name != "ACTIVE":
time.sleep(5)
video = client.files.get(name=video.name)
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "video", "uri": video.uri, "mime_type": video.mime_type},
{"type": "text", "text": (
"This is a 45-minute product meeting. Return:\n"
"1. Decisions made, one bullet each, with the MM:SS where each was agreed\n"
"2. Action items with owners, if named\n"
"3. Open questions nobody answered"
)},
],
)
print(interaction.output_text)
You can also pass a YouTube URL directly as {"type": "video", "uri": "https://www.youtube.com/watch?v=..."}.
What makes video prompts work:
- Use timestamps. Refer to moments as
MM:SS("what's on screen at 12:40?") and ask for timestamps in the answer so you can check them. - Ask about sound and picture together. "Where does what the presenter says contradict the slide on screen?" plays to Gemini's strength.
- Know the sampling rate. By default the model sees 1 frame per second. Fast on-screen changes (a flashing error message, a quick cursor click) can fall between frames, so ask about what's said as well as what's shown.
- Know the limits. With a 1M window, Gemini handles up to about 3 hours of video at low resolution or 1 hour at high resolution.
Images, audio and PDFs
The same input list takes other file types. Images can be inline:
import base64
with open("error-screenshot.png", "rb") as f:
image_bytes = f.read()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": (
"This is a production error. Identify: (1) the error type, "
"(2) the likely root cause from the stack trace, (3) the line most likely "
"responsible, (4) a concrete fix. Don't say 'check your configuration.'"
)},
{"type": "image", "data": base64.b64encode(image_bytes).decode("utf-8"), "mime_type": "image/png"},
],
)
Audio works like video: upload, then pass {"type": "audio", "uri": file.uri, "mime_type": file.mime_type}.
The rule for every modality is the same: say exactly what to extract. "Analyze this call" gets a generic summary. "List every objection the customer raised, quote their words, and categorize each as price, timing or product fit" gets something you can use.
Prompting for long context
A million tokens only helps if the model can find what matters in them.
Label every document. Use clear delimiters and tell the model what each part is:
Below are three research papers to synthesize. Each is labeled with its source.
<paper id="1" source="Stanford 2024 working-memory study">
[paper 1 content]
</paper>
<paper id="2" source="MIT 2025 replication">
[paper 2 content]
</paper>
<paper id="3" source="Journal of Cognitive Science meta-analysis">
[paper 3 content]
</paper>
TASK: Compare the three methodologies and identify where the findings agree
and conflict. Focus on sample size and measurement differences. Cite papers by id.
Put the task after the material. With very long inputs, restating the question at the end, right before the model answers, keeps it from getting lost.
Say where to look. For a codebase: "The repository tree comes first, then file contents. I need every API route with its HTTP method, and every database query with its table." A precise target beats a bigger window.
Ask for citations. "Cite the paper id for every claim" makes gaps and inventions visible.
Grounding with Google Search
Grounding lets Gemini search Google and cite what it finds:
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="What changed in the EU AI Act obligations that took effect in 2026?",
tools=[{"type": "google_search"}],
)
print(interaction.output_text)
The response text carries url_citation annotations with start and end positions, so you can show sources next to the sentences they support.
Use grounding for: current events, prices, product specs, recent releases, and fact-checking.
Skip it for: timeless questions (math, programming concepts), creative tasks, and latency-sensitive routes where the search round trip hurts.
Code execution
For anything that needs a precise computed answer, let Gemini write and run Python:
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=(
"Average order value is $85 with a standard deviation of $42. Assuming a "
"normal distribution, what percentage of orders fall between $50 and $120?"
),
tools=[{"type": "code_execution"}],
)
The sandbox runs for up to 30 seconds per execution and comes with 40+ libraries preinstalled, including pandas, numpy, scikit-learn and matplotlib. You can't install your own, and matplotlib is the only supported charting library.
Common mistakes with Gemini
Using old SDK code. google.generativeai and GenerativeModel examples are everywhere. Current code uses from google import genai and a client.
Filling the window without structure. Label documents, say where to look, and restate the task at the end.
Vague multimodal prompts. "Analyze this video" gets a generic description. Name the things you want extracted, with timestamps.
Leaving Flash-Lite at minimal thinking on reasoning tasks. It's the default and it's right for bulk classification, but it'll underperform on anything multi-step.
Forgetting the price change. Gemini 3.8 Flash's input and output prices double on January 1, 2027. If you're forecasting costs, use the post-promo rates.