If you point an existing Claude Opus 5 integration at claude-opus-5-5 and change nothing else, two kinds of things break. Some requests fail outright with a 400. Others succeed and quietly behave differently: slower turns, longer outputs, an agent that goes silent between tool calls. Prompting Claude Opus 5.5 well is mostly about knowing which is which.
The model itself is good. Anthropic positions it for long-running agentic coding and knowledge work, at $4 and $20 per million input and output tokens. But most of the lessons below are about your harness and your prompts, not the model. Here's what changed, what to delete from your prompts, and what to test.
The four requests that now return a 400
Check these before you touch any prompt copy. They're hard failures.
Thinking can't be disabled. thinking: {"type": "disabled"} returns a 400, and so does the old thinking: {"type": "enabled", "budget_tokens": N} pattern. Omit the field or send {"type": "adaptive"}, which is equivalent. The control for how much the model thinks is now output_config.effort.
Forced tool use is gone. tool_choice: {"type": "any"} and {"type": "tool", "name": "..."} both fail. auto and none still work. More on the replacement in the tool use section.
Thinking blocks are tied to the model and the conversation. Opus 5.5 can read thinking blocks from Opus 5 and earlier Opus, Sonnet and Haiku models, but not from Fable or Mythos models. If you swap models mid-conversation, the blocks it can't read are dropped before the model sees them (they aren't billed). If you edit the system prompt or tools between turns, accounts created on or after August 31, 2026 get a 400 when a block is replayed after such a change. Keep conversations append-only.
computer_20251124 is rejected on the Claude API and Google Cloud. Move to the computer_toolset_20260801 toolset. (Bedrock still accepts the older tool.)
Here's a minimal request that passes all four:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=64000,
thinking={"type": "adaptive"},
output_config={"effort": "medium"},
tools=tools,
tool_choice={"type": "auto"},
messages=messages,
)
# Don't assume content[0] is text. Select by type.
text = "".join(b.text for b in response.content if b.type == "text")
That last line matters. Every response can now begin with a thinking block, and its thinking field is empty by default (display: "omitted"). Code that reads response.content[0].text will break or return nothing.
Calibrate effort instead of inheriting it
Effort is the main dial on Opus 5.5, and because thinking is always on, it's the first thing to adjust when you trade intelligence against latency and cost. Four rules from Anthropic's guidance:
- Set it explicitly. The default is
medium; Opus 5's washigh. A request with noeffortnow runs lighter than it used to. - Don't port your Opus 5 value. Level names don't map to the same amount of thinking across models. Anthropic reports that Opus 5.5 at
mediummatches or beats Opus 5 athighon its coding and knowledge-work evaluations, and thatlowcomes close on several coding evals at much lower cost. Those are Anthropic's numbers on Anthropic's evals, so run your own sweep. - Expect more thinking per turn at the same level, especially at
xhighandmax. If you keep an old setting, expect longer turns and more output tokens. - Leave room in
max_tokens. Thinking counts toward the limit even when you don't see it. A cap sized for Opus 5 with thinking off can truncate replies. For long agentic turns, Anthropic has had good results with the 128,000 maximum.
| If you see | Try |
|---|---|
| Turns slower or pricier than on Opus 5 | Lower effort first; it cuts thinking more reliably than prompt instructions |
| Quality dropped after lowering | Move up one level and re-test |
| Truncated replies | Raise max_tokens |
You want to reserve xhigh / max | Only where you've measured a quality gain |
Changing the top-level effort between requests invalidates the prompt cache. If you need different effort on individual turns, the per-message effort change (beta) keeps the cache intact.
The full parameter reference across vendors is in our reasoning effort controls cheat sheet. For the longer argument on why effort beats clever prompting, see prompting reasoning models.
Stop telling it to think
This is the biggest prompt-level habit to unlearn, and it connects to the chain-of-thought lesson. On a model that reasons by default, "think step by step" and "think carefully before answering" are at best redundant.
Anthropic's guidance for chat applications is specific: if your system prompt tells Claude to think carefully before answering, consider removing it. In their chat testing, removing such a line made replies start sooner with no clear drop in quality. The model decides how much to think; effort is your control.
Two related cleanups:
- Don't ask it to write out its reasoning in the reply as a substitute for thinking. Such prompts may be declined with a
reasoning_extractionrefusal. Setthinking.displayto"summarized"and read the reasoning from the thinking blocks instead. You can still ask for a short explanation of the answer or a summary of what it did. - Re-test any "don't think" workarounds. If your Opus 5 integration ran with thinking disabled, start at
loweffort and measure on your own traffic. If time to first token still matters, a line such as "Answer directly without deliberating" can trim thinking further, but check quality when you add it.
There's one multi-turn quirk worth knowing. In longer chats the model sometimes re-examines an earlier answer while thinking about a new message, even a short follow-up, which adds latency. If you'd rather it treat earlier answers as settled, a two-sentence instruction at the end of the system prompt does that (Anthropic's guide has suggested wording). Skip it for long analyses and agentic work, where a later step may legitimately expose an earlier mistake.
Tool use without forced tool_choice
If your agent used tool_choice: {"type": "any"} to guarantee a tool call or to get structured JSON back, you need a replacement. There are three, depending on why you forced it:
- You wanted valid arguments. Keep
tool_choice: autoand setstrict: trueon the tool schema. - You only wanted JSON. Skip the tool and use structured outputs.
- You wanted the model to call a tool instead of answering in text. Say so in the prompt: name the tool and state when it applies.
That third one is plain instruction-writing. Compare:
Weak: You have access to a search tool.
Better: For any question about current prices, call search_prices first.
Don't answer price questions from memory.
Beyond tool_choice, the general advice from our function calling lesson still holds: precise tool descriptions, narrow parameters, clear error messages. What's new is that the model is also more eager to get going. On loosely specified, multi-app tasks (email, documents, spreadsheets, CRM), Anthropic notes that the information a task depends on often sits somewhere the request doesn't mention, and that one sentence in the system prompt telling the model to explore the relevant sources before acting improved completion in their testing, at the cost of slightly more tool calls and tokens. Because that instruction tells the model to act on what it finds, keep untrusted content out of the records it searches.
Progress updates and unattended agents
Two harness behaviors catch teams off guard, and neither throws an error.
Your UI may go quiet. On Opus 5.5, the short notes the model writes between tool calls arrive as progress-update thinking blocks, not text blocks, and their text is empty at the default display setting. A client that renders only text blocks looks frozen during a long agentic turn. Setting display: "updates" returns a summary of each note, but that's a beta feature behind a header, so check the current docs before depending on it.
Your loop may stop early. On long, multi-part tasks the model keeps the user updated, and sometimes a status update ends the turn with text and no tool call (stop_reason: "end_turn"). A loop that treats every end_turn as "done" will quit partway through. The fix is on the harness side:
- Treat a text-only end of turn as a report, not proof of completion.
- Keep the task's parts in a checklist the model updates (a to-do tool or a file).
- If items are still open and no blocker was stated, send a short user message naming them.
- Cap automatic continuations at two or three so a genuinely stuck run ends and gets reviewed.
You can also reduce early stops with a system prompt addition that names the specific stops you don't want, such as ending a turn with a summary that announces the next step instead of taking it. Add it from the first request, since changing the system prompt mid-session invalidates earlier thinking blocks. And keep your own confirmation step for risky or irreversible actions; skip the addition entirely in human-in-the-loop apps.
Safety classifiers and pasted text
Opus 5.5 runs classifiers for biology, cybersecurity and reasoning extraction. A decline comes back as a normal HTTP 200 with stop_reason: "refusal" and a stop_details object naming the category. Handle it in code, and configure a fallback model rather than letting the user see a blank response.
On the prompt-injection side, Anthropic reports the model resists instructions arriving through tool results, web pages and on-screen content better than any earlier Opus. For text a user pastes into their message, it recommends wrapping each pasted block in matching tags carrying the same short random ID, with a system prompt note explaining the convention. It's one guardrail, not a wall; see our guide to prompt injection in tool-using agents for the rest of the defenses.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
400 on thinking | Disabled or budget_tokens set | Omit thinking or use adaptive; tune effort |
400 on tool_choice | any or named tool | auto plus strict: true, or structured outputs |
Empty text / content[0] errors | Leading thinking block | Select blocks by type |
| Replies cut off | max_tokens too small for thinking | Raise it |
| Turns slower than on Opus 5 | Old effort value, more thinking per turn | Re-run your effort sweep |
| UI silent between tool calls | Progress notes in thinking blocks | Set display, or handle the block type |
| Agent stops mid-task | end_turn read as done | Checklist plus continuation message |
| 400 after editing prompt/tools | Replayed thinking block, changed prefix | Append-only conversation; use mid-conversation system messages |
Migration checklist from Opus 5
- Change the model ID to
claude-opus-5-5. - Remove
thinkingdisabled andenabledsettings; choose aneffortexplicitly. - Replace
tool_choiceanyandtoolwithautoplus a prompt instruction andstrict: true. - Parse responses by block
type. - Raise
max_tokensand re-run an effort sweep on your own evals. - Decide how your UI shows between-tool-call updates.
- Add handling for
stop_reason: "refusal"and forend_turnin agent loops. - Delete "think carefully" lines and any prompt that asks for reasoning to be written out in the reply.
- If you use computer use on the Claude API or Google Cloud, move to the toolset.
If you're coming from older Claude versions or other vendors, start with migrating prompts to new models, and read the extended thinking guide for the thinking model in more depth. For choosing between Opus 5.5, Sonnet and the rest on cost, see which AI model should I use. When your prompts start sprawling, context engineering is the discipline that keeps them manageable, and the prompt library has copy-paste starting points.
One last note on sources: everything above about parameters, defaults and errors comes from Anthropic's own Opus 5.5 documentation as of October 8, 2026. Beta features and defaults move, so check the docs before you ship.



