When OpenAI released o1, I spent the first few hours prompting it the way I'd prompt GPT-4. My results got worse. The model was already reasoning, and my carefully written "first do X, then do Y" instructions were getting in its way.
Two years later that lesson applies to nearly everything. Reasoning models used to be a special tier you reached for on hard problems. In 2026 almost every frontier model reasons before it answers, so prompting reasoning models is just prompting.
Here's what changes in your prompts, and the one setting that matters more than any phrasing.
What's different about reasoning models
A standard model writes its answer token by token, and the answer is the generation. A reasoning model does two passes: a hidden reasoning phase where it works through the problem, then the visible answer.
That's now the default across the big three:
| Reasoning models | Depth control | Can you see the reasoning? | |
|---|---|---|---|
| OpenAI | GPT-6 Astra (recommended for most reasoning work) | reasoning.effort | Summary only |
| Anthropic | Claude Fable 5.1, Opus 5.5, Sonnet 5 | output_config.effort | Summary only |
| Gemini 3.x | thinking_level | Summary only |
The practical effects:
- Better on hard, multi-step problems. Math, planning, debugging, careful analysis.
- Hidden cost. Reasoning tokens are billed as output tokens even though you never see them.
- Less sensitive to prompting tricks. You don't need to coax reasoning out of them. You need to point it at the right target.
What to stop doing
Stop saying "think step by step." It was the most useful phrase in prompting in 2023. On a reasoning model it's redundant: the thinking already happened before the first visible token.
Stop scripting the reasoning. This used to be good practice:
First, identify the key factors. Then, evaluate each factor.
Finally, synthesize a recommendation.
OpenAI's reasoning guide now says the opposite: give the model "the task, constraints, and desired output format" and avoid prescribing intermediate steps. Your script replaces the model's plan with yours, and the model's is usually better.
Stop asking for <thinking> blocks. Asking a reasoning model to write its reasoning into visible tags gets you a second, performed version of the reasoning, and you pay for it on top of the hidden one. If you need to audit the logic, ask for a short rationale or turn on reasoning summaries.
Stop using few-shot examples to teach reasoning. Examples are still great for showing output format and tone. Worked examples that demonstrate how to think are for models that don't think on their own.
What to start doing
Write complete problem specifications. Everything the model needs to reason about has to be in the prompt: constraints, edge cases, what "done" looks like.
Bad: "Write a function to sort a list."
Better: "Write a Python function that sorts a list of integers in ascending
order. It must handle empty lists, duplicates, negative numbers, and lists
of 100M+ elements efficiently. Return a new list rather than sorting in
place. Include type hints and a brief docstring."
Describe the result, not the route.
Bad: "First, analyze the requirements. Second, design the architecture.
Third, identify potential issues."
Better: "Design the database schema for a multi-tenant SaaS application
with these requirements: [requirements]. I need the final schema, the key
design decisions you made, and the main tradeoffs."
Specify the output format precisely. Reasoning models can produce long answers. Say exactly what you want back:
Return a JSON object with these keys:
- "recommendation": one sentence
- "confidence": "high" | "medium" | "low"
- "key_risks": array of strings (max 3)
- "rationale": 2-3 sentences
For anything a program will parse, use the vendor's structured outputs feature rather than relying on the prompt alone.
Ask it to check its own work. "Verify the totals add up before answering" or "confirm every claim is supported by the attached document" still pays off.
Effort is the new model switch
The biggest change since o1 is where the cost-quality trade-off lives. It used to be a model choice: standard model or reasoning model. Now it's mostly an effort setting on the same model.
Higher effort means more hidden reasoning: better answers on hard problems, more latency, more output tokens billed. Lower effort means faster and cheaper, and it's often plenty for classification, extraction and simple Q&A.
OpenAI (Responses API):
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6-astra",
reasoning={"effort": "high", "summary": "auto"},
input="Design a database schema for a multi-tenant SaaS app. [requirements]",
)
print(response.output_text)
Effort values run from none up to max, and which ones are available depends on the model. OpenAI recommends the Responses API over Chat Completions for reasoning models.
Anthropic (Claude):
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "Design a database schema for... [requirements]"}],
)
Claude uses adaptive thinking: the model decides how much to think, and effort (low, medium, high, xhigh, max) steers it. The old thinking={"type": "enabled", "budget_tokens": N} pattern returns an error on current Claude models, so update any code that still uses it. Two Opus 5.5 details matter here: thinking can't be switched off, and effort defaults to medium. Set it explicitly.
Google (Gemini):
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Design a database schema for... [requirements]",
generation_config={"thinking_level": "high"},
)
Gemini 3.x uses thinking_level (low, medium, high, plus minimal on some models). Most models default to medium.
Picking an effort level
| Task | Effort |
|---|---|
| Classification, extraction, routing, simple Q&A | Low |
| Everyday writing, summarization, standard coding | Medium |
| Multi-step debugging, architecture, careful analysis | High |
| Long-running agentic work, hardest problems where correctness beats cost | Highest available (xhigh / max) |
Don't guess. Run the same 20–30 real tasks at two effort levels, compare quality and cost, and pick the lowest level that holds up. Our Promptfoo tutorial shows how to run that comparison in one command.
When a reasoning model is overkill
Even with effort turned down, a flagship reasoning model isn't always the right tool:
- High-volume, latency-critical routes (autocomplete, real-time classification): use the fast tier, like GPT-6 Luna, Claude Haiku 4.5 or Gemini 3.5 Flash-Lite
- Tasks a cheaper model already passes: if your eval says the small model is fine, the big one is burning money
- Creative writing where reasoning isn't the bottleneck
Try the cheapest model at low effort first. Move up only when your tests show it failing. Current prices for every tier are on our models page.
The bottom line
Reasoning models turned prompting from "make the model think" into "tell the model exactly what you need, then let it think." The adjustments:
- Drop "think step by step" and scripted reasoning steps
- Write complete problem specs with constraints and edge cases
- Specify the output format precisely
- Control depth with the effort setting, and measure before you raise it
- Ask for a short rationale or summary when you need to audit the logic
For how this changed classic chain-of-thought prompting, see the updated chain of thought lesson.



