A prompt that works great on one model often underperforms on the other. Not because one model is better — because they're different, and the differences are consistent and learnable.
This is Claude (Fable 5.1, Opus 5.5, Sonnet 5) versus OpenAI's GPT-6 family (Astra, Sol, Luna). Most of what follows compares mid-tier defaults — Claude Sonnet 5 and GPT-6 Sol — since that's where most production traffic lives. The techniques carry across tiers.
The headline differences
Both are capable of most tasks, and for simple prompts the differences are often minor. The gap shows up on something complex, long, or nuanced.
Claude tends to be better at:
- Following nuanced, multi-part instructions precisely
- Long document analysis, with a 1M-token context window across current models
- Tasks where you want explicit structure via XML tags
- Nuanced reasoning that requires integrating multiple considerations
GPT-6 tends to be better at:
- Native structured-output schema validation with strict guarantees
- Vision tasks with a straightforward
input_imagecontent type - Fast, cheap classification and extraction on the Luna tier
- Tool calling with parallel function calls in a well-documented shape
Both families reason internally by default now. What used to be a model-family choice (a fast model vs. a separate reasoning model) is a request-level effort setting on both. See how to prompt reasoning models.
Structural formatting: XML vs. markdown
This is still the biggest practical difference for prompt engineering.
Claude: XML is your friend
Claude was trained with XML tags as a first-class structuring tool:
<context>
You are helping a user write code for a web scraping project.
</context>
<instructions>
Review the code below. Identify:
1. Any bugs that would cause it to fail
2. Performance issues for large-scale use
3. Missing error handling
</instructions>
<code>
[paste code here]
</code>
<output_format>
Use headers for each category. Be specific about line numbers.
</output_format>
Claude reads this as clearly structured information — the tags help it separate context, instructions, data and format specification.
GPT-6: markdown or XML both work
# Context
You are helping a user write code for a web scraping project.
## Task
Review the code below and identify:
1. Any bugs that would cause it to fail
2. Performance issues for large-scale use
3. Missing error handling
## Code
[paste code here]
## Format
Use headers for each category. Be specific about line numbers.
Neither approach is wrong for either model, but these are the native "dialects" each was trained on most heavily. XML is close to mandatory for getting the most out of Claude; markdown is a strong default for GPT-6 but XML works fine there too.
Instruction following: precision vs. flexibility
Claude follows instructions very precisely — sometimes almost literally. If you say "respond in bullet points," it will respond in bullet points even when a paragraph would read better. If you say "do not mention X," it will almost never mention X.
This is good when you want exact compliance. It can be frustrating when you've written slightly imprecise instructions and Claude interprets them literally in a way you didn't intend.
GPT-6 uses more judgment about when instructions should be followed rigidly versus when to deviate for better results. It's more likely to use prose when bullets would be awkward, even if you asked for bullets.
Practical implication: with Claude, write instructions you actually want followed precisely. With GPT-6, you can be slightly looser and the model will exercise judgment — but that also means less predictable formatting behavior.
Long context: Claude and GPT-6 are now close
For long documents — legal contracts, research papers, full codebases, lengthy chat histories — both families now offer roughly a million tokens of context: Claude's current models (Fable 5.1, Opus 5.5, Opus 5, Sonnet 5) and GPT-6 all sit around 1M input tokens. Context size alone rarely decides the choice anymore.
What still differs is structure and mid-document recall. Claude's XML-tag training tends to help it stay oriented across a long, labeled document. For either model, label your sections and restate the task near the end of a very long prompt rather than only at the start.
Tone and personality
Claude tends to be slightly more verbose by default. It hedges, adds caveats, and provides context. This is often good — but sometimes you just want the direct answer.
For Claude, be explicit about brevity:
Answer in one sentence. Do not add caveats.
GPT-6 is slightly more direct by default and tends toward shorter responses. For GPT-6, be explicit when you want depth:
Provide a thorough answer with examples and explanations.
Both models follow these instructions well once you give them. These are tendencies from everyday use, not a controlled study — calibrate against your own prompts rather than taking them as fixed rules.
JSON and structured output
GPT-6 with structured outputs (Responses API):
from pydantic import BaseModel
from openai import OpenAI
client = OpenAI()
class Summary(BaseModel):
summary: str
sentiment: str # "positive" | "negative" | "neutral"
confidence: float
response = client.responses.parse(
model="gpt-6-sol",
input="[text to analyze]",
text_format=Summary,
)
result = response.output_parsed # guaranteed to match the schema
Claude for structured output:
<output_format>
Return a JSON object with exactly these keys:
- "summary": string (1-2 sentences)
- "sentiment": "positive" | "negative" | "neutral"
- "confidence": number between 0 and 1
Return only the JSON object, no other text.
</output_format>
Claude follows structured-output instructions reliably but doesn't have the same native schema-enforcement mode as GPT-6's text_format/output_parsed path. For applications where a malformed field is unacceptable, GPT-6's structured outputs give you a stronger guarantee.
Quick reference
| Task | Recommended starting point |
|---|---|
| Long document analysis | Either (both ~1M context; test structure) |
| Complex, multi-part instructions | Claude |
| Precise formatting compliance | Claude |
| Broad coding assistance | Either (test both) |
| Vision tasks | Either (both take images) |
| Guaranteed schema-valid JSON | GPT-6 (text_format / structured outputs) |
| Creative open-ended writing | Either (preference varies) |
| XML-structured prompts | Claude |
| High-volume, low-cost classification | GPT-6 Luna or Claude Haiku 4.5 |
For most tasks, prompt quality matters more than model choice. A great prompt on GPT-6 usually beats a mediocre prompt on Claude, and vice versa.
When you're optimizing for production: test both models with your actual prompts on your actual use case. Our Promptfoo tutorial sets up a side-by-side comparison in one config file. The differences above are tendencies, not guarantees.
The techniques on MasterPrompting.net — chain-of-thought's 2026 update, few-shot examples, XML structure — apply to both. The differences here are in calibration, not fundamentals. Full specs and prices for both families are on the models page.



