This is a lookup table, not a tutorial. If you want the reasoning behind these settings, why "think step by step" stopped working, how to prompt around adaptive thinking, when to raise or lower effort: that's the full guide. Bookmark this one for the five seconds you spend mid-integration trying to remember whether it's thinking_level or thinking.level.
Three vendors, three parameter shapes, none of them compatible with each other. Here's what to pass, where to put it, and what changes at each setting.
The parameter, at a glance
| Vendor | Field | Where it lives | Values (model-dependent) |
|---|---|---|---|
| OpenAI (Responses API) | reasoning.effort | Top-level reasoning object | none · minimal · low · medium · high · xhigh · max |
| OpenAI (Chat Completions) | reasoning_effort | Top-level field | Same set, model-dependent |
| Anthropic (Claude) | output_config.effort + thinking.type | output_config object, plus thinking: {type: "adaptive"} | low · medium · high · xhigh · max |
| Google (Gemini) | thinking_level | Nested inside generation_config.thinking | minimal · low · medium · high |
No two vendors use the same field name, the same nesting, or the same default. Don't assume you can copy a value across providers: "high" means something different in each column below.
OpenAI: reasoning.effort
response = client.responses.create(
model="gpt-6-astra",
reasoning={"effort": "high", "summary": "auto"},
input="Design a database migration plan for splitting a monolith table.",
)
print(response.output_text)
| Value | Trade-off | Typical use |
|---|---|---|
none | No reasoning phase (not supported on every model) | Simple lookups, formatting |
minimal | Fastest, cheapest reasoning pass | High-volume, low-complexity |
low | Speed and cost optimized | Data analysis, drafting, execution-oriented coding |
medium | Balanced: the default on current models | Agentic coding, research, everyday tasks |
high | Quality prioritized over speed/cost | Complex debugging, deep planning |
xhigh | Extended reasoning | Security review, high-stakes research |
max | Maximum reasoning depth | The hardest problems you have |
Notes:
- Not every model supports every value:
nonein particular is rejected on some current models (e.g. GPT-6 Astra returns a 400 onnone; uselowinstead). Check the specific model's docs before assuming the full range is available. reasoning.summary(values likeauto,concise,detailed) gets you a readable summary of the hidden reasoning. The raw reasoning trace is never returned by the API, regardless of this setting.- OpenAI recommends the Responses API over Chat Completions for reasoning models: it manages reasoning-item state across turns in a way Chat Completions isn't built for.
- Reasoning tokens are billed as output tokens even though you never see them.
Anthropic (Claude): thinking + output_config.effort
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "Design a database migration plan..."}],
)
| Model | Default effort | Can thinking be disabled? | budget_tokens |
|---|---|---|---|
| Claude Fable 5.1 | high | No ({type: "disabled"} returns 400 | Removed) 400 |
| Claude Opus 5.5 | medium (one step below Opus 5's default (set it explicitly) | No) 400 at every effort level | Removed: 400 |
| Claude Opus 5 | high | Only at effort high or below | Removed: 400 |
| Claude Sonnet 5 | high | Yes ({type: "disabled"} accepted | Removed) 400 |
| Claude Haiku 4.5 | n/a (no adaptive thinking) | Yes, via omission | Still required for thinking: the one current-model exception |
| Effort value | Trade-off |
|---|---|
low | Cheapest, fastest: good for sub-agents and simple tasks |
medium | Balanced |
high | Minimum recommended for intelligence-sensitive work |
xhigh | Best setting for most coding/agentic work on current models |
max | Correctness over cost |
Notes:
budget_tokens(the oldthinking: {type: "enabled", budget_tokens: N}pattern) returns a 400 on every current Claude model except Haiku 4.5, which still requires it for thinking.- Thinking display defaults to
"omitted"(empty thinking text) on current models: setdisplay: "summarized"explicitly if you want a readable reasoning summary streamed back. - Assistant message prefill is rejected (400) on all current Claude models, this interacts with reasoning because prefill was a common way to force output shape; use
output_config.format(structured outputs) instead. - Forced
tool_choice("any"or a named tool) returns a 400 specifically on Claude Fable 5.1 and Claude Opus 5.5: not on Sonnet 5 or Opus 5.
For the full behavioral picture on Claude specifically (adaptive thinking, effort tuning by task type, prompting patterns) see the Claude extended thinking guide.
Google (Gemini): thinking_level
response = client.interactions.create(
model="gemini-3.8-flash",
input="Design a database migration plan for splitting a monolith table.",
generation_config={"thinking_level": "high"},
)
| Value | Trade-off |
|---|---|
minimal | Fewest possible thinking tokens: low-complexity tasks only, not supported on every model |
low | Fast, low-latency responses |
medium | Balanced speed/reasoning: the default on most Gemini 3.x models |
high | Dynamic thinking for complex reasoning |
Notes:
minimalsupport varies by model: some Flash-tier models default to it, while some newer Flash models (e.g. Gemini 3.8 Flash) reject it outright. Confirm support for your specific model before relying on it.- The older numeric
thinking_budgetparameter still exists on some models but is no longer the recommended control: usethinking_levelin new integrations. thinking_summaries(set to"auto") returns a readable summary of the reasoning. Raw thoughts are not directly exposed.max_output_tokenscovers thought tokens too, if output is getting truncated, lowerthinking_levelrather than trying to carve out headroom with a token limit.
Cross-vendor cheat sheet
| Question | OpenAI | Claude | Gemini |
|---|---|---|---|
| Field name | reasoning.effort | output_config.effort | thinking_level |
| Lowest value | none (model-dependent) | low | minimal (model-dependent) |
| Highest value | max | max | high |
| Default (typical current model) | medium | high (medium on Opus 5.5) | medium |
| Can reasoning be fully disabled? | Yes, on models supporting none | Only on Sonnet 5 among current models | Effectively, via minimal on supporting models |
| Get a reasoning summary | reasoning.summary | thinking.display: "summarized" | thinking_summaries: "auto" |
| Raw reasoning ever exposed | No | No | No |
| Billed as output tokens | Yes | Yes | Yes |
Rules of thumb that hold across all three
- Higher effort costs more and is slower, but "cost" here means both token spend and wall-clock latency, since reasoning happens before the first visible token.
- None of the three expose raw reasoning. Every "see the thinking" feature is a summary, generated separately from the actual reasoning trace. Don't build a feature that depends on inspecting the real chain of thought: it isn't available on any of these three vendors as of this writing.
- Defaults are not uniform even within one vendor's lineup. Claude Opus 5.5 defaults to
mediumwhile Opus 5 defaults tohigh: same vendor, same effort scale, different default. Set effort explicitly per route rather than trusting an inherited default when you're not certain which exact model will serve the request. - Lower effort isn't just "worse and cheaper." For routine tasks (classification, simple extraction, high-volume chat) a lower effort setting on a current-generation model often matches or beats a higher effort setting on the previous generation, at a fraction of the cost. Measure before assuming you need
higheverywhere.
If you're mid-migration from an older model generation and want the full breaking-changes list (deprecated parameters, prefill removal, forced tool_choice errors) that's covered in migrating prompts to new models. For the prompting-technique side of reasoning models rather than the parameter reference, go to the full guide.



