A reasoning model is one that works through a problem internally before it produces a visible answer. That internal pass, often called "thinking" or a "reasoning trace," isn't part of the response you see. It's a separate phase the model runs first, and you can usually control how long or thorough that phase is with a setting: effort, thinking level, or something similarly named depending on the vendor.
This is a newer default than it might feel like. Models like GPT-4o didn't reason internally at all. If you wanted step-by-step reasoning out of GPT-4o, you had to prompt for it explicitly, and a separate model line (o1, o3) existed specifically to do the reasoning that GPT-4o couldn't. As of today, current flagship-and-above models from OpenAI, Anthropic, and Google all reason internally by default. The separate "reasoning model" category has effectively merged into the main model line.
How the current controls work
Each vendor exposes this differently, and the parameter names don't map cleanly across providers:
- OpenAI:
reasoning.effort, with values fromnone(not supported on every model) throughminimal,low,medium,high,xhigh, andmax. - Anthropic:
thinking: {type: "adaptive"}combined withoutput_config.effort, using a similar low-to-max scale. - Google:
thinking_level, nested inside the generation config.
Higher effort means the model spends more time and tokens reasoning before it answers, which typically improves quality on genuinely hard problems and adds latency and cost on everything else. The full parameter-by-parameter breakdown, including exact field names and nesting for each vendor, lives in the reasoning effort controls reference. Keep that open in a tab when you're integrating; the exact field names aren't worth memorizing.
For the reasoning behind these settings, including why effort tuning matters and how it interacts with prompt design, see the full prompting reasoning models guide. This lesson covers the version you need to start prompting reasoning models correctly today.
Why "think step by step" stopped being necessary
For years, one of the most reliable prompting techniques was telling the model to reason through a problem explicitly: "think step by step," "explain your reasoning before answering," or providing a scripted sequence of steps to follow. That advice made sense for models that answered immediately, token by token, with no internal deliberation. Prompting them to externalize a reasoning process before committing to an answer measurably improved accuracy.
Current reasoning models already do that deliberation internally, automatically, on every call where reasoning is enabled. Telling a reasoning model to "think step by step" is redundant at best. At worst, it can push the model toward a rigid, scripted process when the internal reasoning phase would have found a better approach on its own.
This doesn't mean structure in your prompt stops mattering. It means the structure that helps has changed. Instead of scripting the model's thought process, you now specify the task clearly, state the constraints it needs to respect, and describe the output format you want. The model figures out how to get there.
How to prompt a reasoning model
Three things matter more than a scripted chain of steps:
State the task precisely. Reasoning models are good at filling in a path from problem to solution, but they still need to know what the actual problem is. Vague tasks get vague reasoning. "Review this function for bugs" is weaker than "review this function for off-by-one errors and unhandled null inputs, and list each issue with the line number."
Give constraints, not instructions. Instead of "first check the types, then check the edge cases, then check performance," tell the model what matters: "this needs to handle null and empty-array inputs correctly, and it runs in a hot path so allocations matter." The model will structure its own reasoning around those constraints, often catching things a scripted checklist would have missed.
Specify the output format explicitly. Reasoning happens in a phase you don't see, so the model needs separate, clear instructions for what the final answer should look like: a JSON object with specific fields, a numbered list of issues, a single paragraph. Don't assume the reasoning phase will also solve your formatting problem.
A prompt for a reasoning model might look like this:
Review this checkout function for correctness issues around
discount code validation. It needs to handle expired codes,
codes that don't apply to items in the cart, and codes used
past their redemption limit.
Return a JSON array of issues, each with `line`, `severity`
(high/medium/low), and `description`.
No numbered steps, no "think carefully about each case." The task, the constraints, and the output shape are all it needs. At a higher effort setting, the model will spend more internal reasoning on edge cases you didn't explicitly list, which is exactly the point of paying for that setting.
When to raise or lower effort
Default to a medium setting for most work; it's what current models default to for a reason. Raise it for genuinely ambiguous or high-stakes tasks: security review, architectural decisions, debugging something that's resisted a first pass. Lower it for high-volume, low-complexity work where speed and cost matter more than squeezing out marginal quality: classification, formatting, simple extraction.
Watch for a common mistake: raising effort to fix a problem that's actually a prompting problem. If the model's answers are wrong because the task was ambiguous or the constraints weren't stated, more reasoning time won't fix that reliably. Fix the prompt first, then tune effort.
That's the core of prompting reasoning models: trust the internal reasoning phase to handle the "how," and spend your prompt on the "what."
Later in this track, you'll get into multimodal models, world models and JEPA, and open versus closed weights, which round out the rest of what's changed in how models work.