Skip to main content
All Model Guides
Model GuideClaudeAnthropicXML tagssystem promptsextended thinking

How to Prompt Claude: Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5

Claude responds uniquely well to XML tags, explained constraints and clear output specs. The prompting patterns, thinking and effort settings, and model choices that work best with Claude in 2026.

8 min readVerified

Claude, built by Anthropic, has distinct strengths and quirks compared to other frontier models. It's particularly good at nuanced instruction-following, long document analysis, and careful reasoning — but it responds best to prompts written in specific ways.


What Makes Claude Different

Trained for instruction-following over raw task completion. Claude is trained with an emphasis on understanding what you actually want, not just what you literally said. This means it handles ambiguity better than most models — but it also means vague prompts lead to vague answers.

Responds exceptionally well to XML tags. Anthropic trained Claude with XML as a structuring convention. Tags like <context>, <instructions>, and <output_format> help Claude parse complex prompts cleanly. This is especially useful for multi-part instructions with multiple types of content.

Long context window. Current Claude models (Fable 5.1, Opus 5.5, Opus 5 and Sonnet 5) take up to 1M tokens of input, and Haiku 4.5 takes 200K. That makes Claude well-suited to document analysis, long-form editing, and large codebase review, as long as you structure the input with tags.

Tendency toward caution. By default, Claude adds safety caveats and hedges where it's uncertain. For professional use cases, you often need to calibrate this via system prompts.


Core Prompting Patterns for Claude

Use XML Tags for Structure

For any prompt with multiple sections, XML tags reduce ambiguity significantly:

<context>
You are reviewing a research proposal for a neuroscience grant committee.
The proposal is for a 3-year study on working memory consolidation.
</context>

<instructions>
Evaluate this proposal on three dimensions: scientific rigor,
feasibility, and novelty. For each, give a score from 1-5 and a brief
justification. Then give an overall recommendation.
</instructions>

<proposal>
[paste proposal here]
</proposal>

<output_format>
## Scientific Rigor: [score]/5
[justification]

## Feasibility: [score]/5
[justification]

## Novelty: [score]/5
[justification]

## Recommendation
[decision + key reason]
</output_format>

The tags serve as named sections that Claude can refer back to as it generates each part of the response.

Write System Prompts That Explain the "Why"

Claude responds notably better when you explain the reasoning behind constraints, not just state them:

<!-- Less effective -->
<instructions>Always recommend consulting a doctor.</instructions>

<!-- More effective -->
<instructions>
Always recommend consulting a doctor for medical questions.
This is because our platform is used by patients who may act on advice
without professional review, and we want to ensure they have appropriate
oversight for any health decisions.
</instructions>

The explanation gives Claude the context to apply the constraint intelligently across edge cases.

Give Claude Permission to Be Direct

By default, Claude adds caveats, hedges, and considers multiple perspectives. If you want directness, ask for it:

You are an expert Python developer reviewing this code. Give me direct,
specific feedback. Don't hedge with "this could potentially be a problem" —
tell me what IS a problem and exactly how to fix it.

Or in a system prompt:

<style>
Be direct and opinionated. If something is wrong, say it's wrong.
If there's a clear best option, recommend it rather than listing pros
and cons of all options. I can handle direct feedback.
</style>

Use Multi-Turn Conversation Strategically

Claude maintains context effectively across long conversations. For complex tasks, breaking them into turns often produces better results than a single massive prompt:

Turn 1: "Here's my business model. Analyze it and identify the three
         biggest risks."

Turn 2: "For risk #2 [competitive saturation], what are the three most
         viable mitigation strategies for a bootstrapped startup?"

Turn 3: "Draft a 1-page risk mitigation plan for the board covering
         what we discussed."

Each turn builds on the previous analysis rather than asking Claude to do everything at once.


Thinking and Effort

Current Claude models reason before answering. On Fable 5.1, Opus 5.5 and Sonnet 5 this is adaptive thinking: Claude decides how much to think for each request, and you steer it with effort.

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    thinking={"type": "adaptive", "display": "summarized"},
    output_config={"effort": "high"},
    messages=[{
        "role": "user",
        "content": "Design the database schema for a multi-tenant SaaS application...",
    }],
)

for block in response.content:
    if block.type == "thinking":
        print("Reasoning summary:", block.thinking)
    elif block.type == "text":
        print("Response:", block.text)

What changed from older tutorials:

  • budget_tokens is gone. thinking={"type": "enabled", "budget_tokens": N} returns a 400 on Fable 5.1, Opus 5.5, Opus 5 and Sonnet 5. Haiku 4.5 is the exception and still uses it.
  • Effort controls depth: low, medium, high, xhigh or max. Use low for simple tasks and sub-agents, high as the minimum for intelligence-sensitive work, and xhigh for most coding and agentic tasks.
  • Fable 5.1 and Opus 5.5 always think. You can't disable it; lower the effort instead. Opus 5.5 also defaults to medium effort rather than high, so set it explicitly.
  • The raw reasoning is never returned. The default display is "omitted" (empty thinking text). Set display: "summarized" for a readable summary.

Don't add "think step by step" or ask for <thinking> tags on these models. The reasoning already happens in the thinking phase. The full walkthrough is in our Claude extended thinking guide.

Prompts written for older Claude models

Two habits that worked on Claude 3.x and 4.x now backfire:

  • Over-prescriptive prompts. Anthropic's own migration guidance notes that prompts written for prior models are often too prescriptive and reduce output quality on the newest ones. If a long list of "always do X, then Y" rules is producing stiff results, cut it back to the goal, the constraints and the output format.
  • Assistant prefill. Starting Claude's reply for it (for example, prefilling { to force JSON) returns a 400 on Fable 5.1, Opus 5.5, Opus 5 and Sonnet 5 (Haiku 4.5 still accepts it). Use structured outputs or a clear format instruction instead.

System Prompt Best Practices

The Minimal Effective System Prompt

A focused 200-token system prompt often outperforms a sprawling 2,000-token one. Every instruction you add is one more thing Claude has to track and potentially conflict with.

Start with the minimum:

<persona>
You are a senior software engineer specializing in Python and system design.
</persona>

<constraints>
- Write production-ready code, not tutorial snippets
- If a requirement is ambiguous, ask before assuming
- Don't add complexity beyond what was asked
</constraints>

Then add instructions only when the output consistently fails in a specific way.

The Claude Cookbook (Proven Patterns)

GoalTechnique
Prevent AI-sounding output"Write in a direct, professional tone without filler phrases"
Stop over-hedging"Give a direct answer; I'll ask if I need qualifications"
Get specific code"Write complete, working code — no # TODO or placeholder comments"
Improve research quality"Distinguish between: established consensus, emerging evidence, and your inference"
Control length"Be concise. If I want elaboration, I'll ask."
Prevent over-engineering"Solve the problem stated, not a hypothetical harder version"

Claude Model Versions

ModelAPI IDContextPrice (in / out per 1M)Best for
Claude Fable 5.1claude-fable-5-11M$10 / $50Anthropic's most capable model; hardest long-running work
Claude Opus 5.5claude-opus-5-51M$4 / $20Demanding coding and agentic work
Claude Opus 5claude-opus-51M$5 / $25General-purpose Opus tier
Claude Sonnet 5claude-sonnet-51M$2 / $10Most production tasks
Claude Haiku 4.5claude-haiku-4-5200K$1 / $5High-volume, simple tasks and sub-agents

Rule of thumb: Start with Sonnet 5. Move to Opus 5.5 when Sonnet gives inconsistent results on coding or agent tasks, and to Fable 5.1 for the hardest problems. Try lowering effort before moving down a tier: a bigger model at lower effort sometimes beats a smaller one at high effort. Use Haiku 4.5 for bulk classification and extraction.

The full cross-vendor table is on the models page.


Common Mistakes With Claude

Being too polite in system prompts. "Claude, please try to be concise if you can" is weaker than "Be concise. Omit preamble and filler." Claude doesn't need diplomatic phrasing — it needs clear instructions.

Not explaining constraints. If you tell Claude not to do something, say why. It helps Claude apply the constraint to edge cases you didn't anticipate.

Fighting Claude's safety defaults with tricks. This wastes time and often produces worse results. Instead, give Claude legitimate context: who you are, what you're building, why you need a specific type of response. Honest instructions work better than jailbreak attempts.

Using one massive prompt when multiple turns would work better. For complex analysis tasks, a conversation often produces better results than a single 10,000-token prompt.

Relying on prefill or forced tool choice. Both return errors on the newest models: prefill on everything except Haiku 4.5, and tool_choice set to any or a specific tool on Fable 5.1 and Opus 5.5. Use auto with a clear instruction, strict: true on the tool, or structured outputs.

Not testing system prompt length. If you have a long system prompt, try halving it and see if the behavior changes meaningfully. Often, half the instructions do 95% of the work.

Want to compare models side by side?

See how OpenAI, Claude, Gemini and open-weight models stack up for different use cases.

View model comparison →