I inherited a prompt library last month that hadn't been touched since GPT-4o was the default model. Half the prompts still opened with "think step by step," one of them prefilled the assistant turn to force JSON, and a tool-calling wrapper hard-locked tool_choice to a specific function on every call. None of that is a style choice anymore: on current models, some of it just breaks.
If you've got a prompt library that's more than a year old, you're not migrating to a slightly-better model. You're migrating to a different default behavior: reasoning is on by default now, several request parameters that used to be routine now return a 400, and prompts that were carefully engineered for a non-reasoning model can quietly make a reasoning model worse. Here's what actually breaks, what changes in behavior even when nothing errors, and a checklist for auditing a prompt library instead of finding out in production.
The starting point: what "old" and "new" mean here
Old, for this post: prompts and integration code built against GPT-4o, Claude 3.x (Opus/Sonnet/Haiku 3), and Gemini 1.5 or 2.0.
New: GPT-6, Claude Fable 5.1 / Opus 5.5 / Sonnet 5, and Gemini 3.x.
The gap between those two generations is bigger than a single version bump because it spans the point where "reasoning model" stopped being a special tier and became the default. GPT-4o didn't reason before answering. GPT-6 does, unless you explicitly turn it down. Same story for current Claude and Gemini. That single shift is behind most of what follows.
What breaks outright (400 errors, not degraded output)
These aren't style suggestions: they're hard failures on current Claude models. If your codebase touches the Anthropic API, check for all three before you touch a single line of prompt copy.
Fixed thinking budgets are gone. The old thinking: {type: "enabled", budget_tokens: N} pattern (where you hand-picked a token ceiling for the model's reasoning) returns a 400 error on every current Claude model except Haiku 4.5. Current models use thinking: {type: "adaptive"} instead, paired with output_config: {effort: "..."} to control depth. Adaptive thinking means Claude decides how much to think per-request; effort is your dial on that, not a hard token cap.
Assistant prefill is gone. If your old code started Claude's response for it (a common trick to force JSON output, e.g. seeding the assistant turn with {) that now returns a 400 on every current Claude model. Anthropic's replacement is structured outputs (output_config: {format: {...}}) or, for looser formatting control, an explicit instruction in the system prompt. If your codebase has a prefill= argument anywhere in an Anthropic client wrapper, it needs to come out.
Forced tool_choice is gone on two specific models. tool_choice: {"type": "any"} or {"type": "tool", "name": "..."} returns a 400 specifically on Claude Fable 5.1 and Claude Opus 5.5: not on Sonnet 5 or Opus 5, which still accept it. If your agent framework force-calls a tool to guarantee structured output, and you're moving to Fable 5.1 or Opus 5.5, swap that for tool_choice: "auto" plus an explicit prompt instruction naming the tool, strict: true on the tool schema to keep arguments valid, or structured outputs if the forced call only existed to get JSON back.
None of this is unique to Anthropic's API design, it's the same shape of change other providers have made as reasoning became default: parameters that assumed you were steering a non-reasoning model around its limitations stop making sense once the model reasons on its own.
What changes in behavior without erroring
This is the harder category, because nothing tells you it happened. Output quality shifts and you have to notice.
"Think step by step" can now hurt. On a non-reasoning model like GPT-4o or Claude 3.5, spelling out an explicit reasoning sequence ("first identify X, then check Y, then compute Z") was often the single highest-leverage prompt technique available. It gave the model a scratchpad it didn't otherwise have. On a reasoning-by-default model, the model already builds its own internal plan before answering. A rigid, prescriptive script you bolt on top doesn't add reasoning, it overrides the model's own plan with a worse one, especially when your script doesn't match how the model would have actually approached the problem. I've seen this cut output quality on ported prompts with no other changes. The fix isn't to delete all structure (clear task definitions, constraints, and output format still matter enormously) it's to stop prescribing the intermediate steps and instead describe the destination. For the deeper version of this argument, see the chain-of-thought lesson and the full reasoning models guide.
Reasoning tokens are billed and invisible. Under GPT-4o, every output token you paid for was output you could show the user. Under GPT-6 or current Claude, a chunk of output tokens go to a hidden reasoning phase you can't see and can't skip: you can only summarize it (reasoning.summary on OpenAI, display: "summarized" on Claude's thinking block, thought summaries on Gemini). If your cost model or your latency budget was built on GPT-4o's token math, it's now wrong in both directions: costs went up because of the hidden reasoning tokens, and perceived latency went up because there's a thinking phase before the first visible token. Budget and UX (loading states, streaming) both need a second look.
Default effort levels differ by model, not just by provider. This one trips people up specifically because it's not uniform even within a single vendor's lineup. On most current Claude models the default effort is high; on Claude Opus 5.5 specifically it's medium. If your migration script blindly ports "no effort parameter set" from an old integration, you'll get inconsistent depth depending on which exact model you land on. Set effort explicitly per route rather than relying on defaults, see the reference table in our reasoning effort controls cheat sheet for exact values across all three vendors.
Structured output and JSON mode moved. If your GPT-4o integration leaned on prefill or a loosely-worded "respond only in JSON" instruction to get parseable output, current models on all three vendors have dedicated structured-output configuration: output_config.format on Anthropic, the equivalent JSON-schema response format on OpenAI's Responses API, structured output mode on Gemini. These are more reliable than prompt-level instructions and, on Claude specifically, are now the only supported way to force response shape since prefill is gone.
The Responses API is the recommended surface on OpenAI, not Chat Completions. Chat Completions still works, but OpenAI recommends the Responses API for reasoning models specifically, it handles the reasoning-item bookkeeping (passing prior reasoning context back on multi-turn calls) in a way Chat Completions wasn't designed for. If your integration is still on Chat Completions purely out of inertia, this migration is a reasonable point to move it.
A practical audit process
Don't try to rewrite the whole library in one pass. Triage first, then work in order of risk.
1. Grep for the hard breaks first. Before touching prompt copy, search your codebase for the patterns that error outright:
budget_tokens(any Anthropic call)- Assistant-role prefill: a hardcoded partial assistant message passed as the start of the response
- Forced
tool_choice("any"or a named tool) on any call that might route to Fable 5.1 or Opus 5.5 - Hardcoded old model ID strings (
gpt-4o,claude-3-...,gemini-1.5-.../gemini-2.0-...) anywhere in config, not just in obvious "model" fields, I've found them in fallback logic, in retry handlers, and in test fixtures that quietly pin an old model for "speed"
Fix these first. They're the ones that will fail loudly and immediately, which paradoxically makes them the easy tickets: you'll know the instant they're wrong.
2. Inventory every prompt by function, not by file. Group prompts into rough categories: classification/extraction (usually low-stakes to migrate), agentic/tool-use (medium risk, touches the tool_choice issue), and long-form generation or multi-step reasoning tasks (highest risk for the "step by step now hurts" problem). This lets you prioritize the audit instead of reading every prompt cold.
3. For each prompt, ask three questions:
- Does it script the reasoning process explicitly ("first do X, then Y, then Z")? If yes, flag for rewrite: describe the goal and constraints, let the model plan.
- Does it rely on prefill, a fixed thinking budget, or forced tool choice anywhere in the surrounding code? If yes, that's a code change, not a prompt change: see the breaking-changes list above.
- Does it depend on a specific effort/reasoning setting that was never made explicit? If yes, make it explicit now: don't inherit a default from a model you're leaving.
4. Re-run evals before you ship, not after. If you don't have an eval set for a given prompt, this migration is a good forcing function to build even a small one: 10-20 representative cases with a clear pass/fail or scoring rubric. The failure mode here isn't usually "the new model is worse," it's "the new model behaves differently in ways your eyeball won't catch across 200 prompts but a scored eval will."
5. Check anything that touches OpenAI's eval or fine-tuning tooling separately. If part of your stack depends on infrastructure that's since been retired or restructured, that's a parallel migration track, see the OpenAI evals shutdown migration post if that applies to you.
6. Roll out by risk tier, not all at once. Move low-stakes classification and extraction prompts first, they're least likely to be sensitive to the reasoning-behavior shift and give you a fast signal on whether your general migration approach (model IDs, auth, SDK version) is sound. Save the complex agentic and long-form prompts, where the step-by-step rewrite risk is highest, for last, after you've validated the mechanical parts of the migration elsewhere.
What I'd tell a team starting this today
Separate the migration into two passes and don't let them blur together. Pass one is mechanical: swap model IDs, remove parameters that now 400, update SDK calls to the current request shape. Pass two is behavioral: audit the actual prompt text for instructions that assumed a non-reasoning model, and re-tune effort settings per route instead of inheriting whatever the old integration happened to set. Teams that skip pass two usually ship a working migration that quietly produces worse output, because nothing errored: it just reasons differently now than the prompt was written to expect.
If you're touching Claude specifically, the extended thinking guide covers the adaptive thinking and effort model in more depth than fits here. If you want the exact parameter names and values across all three vendors side by side while you're doing the audit, keep our reasoning effort controls reference open in another tab: it's built as a lookup table, not a read-through.
Check our model comparison page for a current snapshot of what each vendor's lineup supports before you commit a migration plan to a specific model, these lineups have moved fast enough in 2026 that a plan written against last quarter's model names is already stale by the time you finish reading it.



