I've now ported the same tool-calling agent to all three major providers more times than I'd like to admit, and every single time I forget one detail, Claude wants all tool_result blocks in a single message, OpenAI wants the arguments as a JSON string you parse yourself, Gemini's newest SDK has renamed half the parameters again. The concepts are identical across vendors. The syntax is not, and the gap between "I know how tool calling works" and "I got this specific SDK call to compile" is where an afternoon disappears.
This site already has posts on structured output basics and JSON mode: read those first if you need the "why does this matter" case for schema-constrained output. This post is narrower and more current: the exact 2026 syntax for structured outputs and tool/function calling on OpenAI's Responses API, Claude's Messages API, and Gemini's google-genai SDK, side by side, plus the handful of places where the concept that transfers cleanly still trips people up in practice.
The concept that's identical everywhere
Strip away the SDK differences and every provider is solving the same two problems with the same two mechanisms:
Structured outputs constrain the model's final text response to match a schema you provide (typically JSON Schema) so you get a parseable object back instead of prose you have to regex out of a paragraph. This is for when you want data, not action.
Tool calling (also called function calling) lets the model, instead of answering directly, emit a request to call one of the functions you've described, with arguments that match a schema. Your code executes the actual function (the model never runs anything itself) and you feed the result back in so the model can use it to continue the conversation or finish the answer. This is for when the model needs information or capability it doesn't have on its own: read the function calling lesson if that distinction is new.
Every provider also now defaults to models that reason internally before responding (see prompting reasoning models) which matters for tool calling specifically: expect multi-turn loops where the model reasons, calls a tool, reasons about the result, and possibly calls another tool, rather than one clean request-response-done cycle. Design your tool-execution loop to run until the model stops requesting tools, not for exactly one round trip.
With that shared model in place, here's where each vendor actually differs.
OpenAI: the Responses API
OpenAI's current surface is the Responses API (the successor to Chat Completions, which still works but isn't where new capability lands first).
Structured outputs use text_format with a Pydantic model (Python) or equivalent schema object, through the .parse() convenience method:
from pydantic import BaseModel
from openai import OpenAI
client = OpenAI()
class Ticket(BaseModel):
category: str
priority: int
escalate: bool
response = client.responses.parse(
model="gpt-6-sol",
input="Classify this ticket: 'Charged twice for my annual plan, need a refund today.'",
text_format=Ticket,
)
ticket = response.output_parsed
output_parsed is the whole point, no manual JSON parsing, no schema-validation step of your own, because the API enforces the schema at generation time ("strict" structured outputs, not best-effort JSON mode).
Tool calling defines tools as flat dicts with type: "function", a JSON Schema parameters block, and strict: true if you want the same enforcement guarantee on the arguments:
tools = [{
"type": "function",
"name": "get_order_status",
"description": "Look up the shipping status of an order by its ID.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
"additionalProperties": False,
},
"strict": True,
}]
response = client.responses.create(model="gpt-6-sol", tools=tools, input=input_items)
You return results as function_call_output items, matched by call_id, appended to the input list for the next call:
import json
for item in response.output:
if item.type == "function_call":
args = json.loads(item.arguments)
result = get_order_status(**args)
input_items.append({
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(result),
})
The two things people get wrong on OpenAI: forgetting arguments comes back as a string you have to json.loads() yourself (not a pre-parsed dict), and forgetting that strict: true requires additionalProperties: False and every property listed in required: the strict mode schema rules are stricter than plain JSON Schema.
Claude: the Messages API
Claude's tool definitions use input_schema instead of parameters, and the request/response shape is block-based rather than item-based.
tools = [{
"name": "get_order_status",
"description": "Look up the shipping status of an order by its ID.",
"input_schema": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
}]
response = client.messages.create(
model="claude-sonnet-5", max_tokens=1024, tools=tools,
messages=[{"role": "user", "content": "Where is order A-1042?"}],
)
Two Claude-specific details that don't have a clean OpenAI or Gemini equivalent:
block.input is already a Python dict: Claude doesn't make you parse a JSON string the way OpenAI does. Easy to get backwards if you're porting code between the two.
Results come back as tool_result content blocks, and every tool_result block from a given turn has to go in a single user message together, not split across multiple messages:
tool_results = [
{
"type": "tool_result",
"tool_use_id": block.id,
"content": get_order_status(**block.input),
}
for block in response.content
if block.type == "tool_use"
]
messages.append({"role": "user", "content": tool_results})
If Claude made two parallel tool calls in one turn, splitting their results into two separate user messages will break, parallel tool use is trained around getting all the results back together, and the API will error or the model will get confused about which result matches which call.
The other current gotcha: on Claude Opus 5.5 and Claude Fable 5.1, forcing tool choice (tool_choice: {"type": "any"} or naming a specific tool) returns a 400 error (type 'tool' and 'any' are not supported for this model). Earlier Claude models supported forced tool choice; these two don't. If your code (or a framework you're using, like some LangChain integrations) forces tool choice to guarantee structured output, you'll need to switch to tool_choice: {"type": "auto"} (the default) combined with strict: true schema validation, plus a prompt instruction telling the model when to use the tool. This is a real breaking change if you're upgrading from Opus 5, check any framework wrapper you depend on before switching models, since several had to patch around it.
For structured JSON output that isn't really "calling a tool" conceptually, Claude also supports a direct output_config parameter for schema-formatted output without going through the tool-use machinery at all, worth reaching for when you want a shaped response but aren't modeling it as a "function the model calls."
Gemini: the current google-genai SDK
Gemini's tool interface is the most different of the three, both in shape and in how broadly "tools" is scoped, Google bundles built-in tools (search grounding, code execution) into the same list as your custom functions.
response = client.interactions.create(
model="gemini-3.8-flash",
input="What's the shipping status of order A-1042?",
tools=[
{"type": "function", "function": get_order_status_schema},
{"type": "google_search"},
],
)
Built-in tools like {"type": "google_search"} for search grounding and {"type": "code_execution"} for sandboxed Python execution sit in the exact same tools array as your custom function definitions: there's no separate "built-in vs. custom" registration path the way there sort of is on OpenAI (with the Assistants/Agents-style hosted tools) or Claude (which has a small, separately-documented set of server-side tools). If you're only used to OpenAI or Claude's tool model, this unified list takes a minute to get used to, but it also means combining structured output, custom function calls, and search grounding in a single request is more natural on Gemini than bolting the same combination onto the other two.
Structured output is configured through the generation config with an explicit output schema rather than a parse()-style convenience wrapper bound to a Pydantic class the way OpenAI's is, Google's Python and JS SDKs both support defining that schema with Pydantic or Zod respectively, but the mechanism is "set a schema in generation config" rather than "call a schema-aware parse method." Worth checking the current google-genai docs directly before shipping, since Google has moved parameter names between SDK versions more aggressively than OpenAI or Anthropic have in 2026, the migration from the older google-generativeai package to google-genai changed several call shapes outright, so a Stack Overflow answer or blog post from more than a few months back may already be stale syntax.
What actually transfers, and what to design around
If you're building something that needs to run on more than one provider, for cost, redundancy, or because different models are better at different parts of your pipeline: a few patterns hold up:
Design your tool schemas in plain JSON Schema first, provider-agnostic, then map them into each SDK's specific wrapper (parameters for OpenAI, input_schema for Claude, the Gemini function dict) at the call site. The schema itself (types, required fields, descriptions) is portable even though the field name holding it isn't.
Don't assume argument parsing is uniform. Write a thin adapter layer that normalizes "arguments as a JSON string" (OpenAI) and "arguments as a parsed dict" (Claude) into one shape before your business logic touches them. This is the single most common silent bug when porting agent code between providers: a TypeError from trying to call .get() on a string, or a JSON parse error from trying to json.loads() a dict.
Don't hardcode forced tool choice as your structured-output strategy if Claude is one of your targets. Prefer strict: true (all three vendors support some flavor of schema-strict tool arguments now) plus a clear system instruction over relying on the API to force a tool call: it's more portable and it's required on the newest Claude models anyway.
Budget for multi-turn tool loops, not single round trips, given all three vendors' current models reason internally and may take several reasoning-then-tool-call cycles, especially with parallel tool calls in a single turn. Your execution loop should keep running (dispatching tool calls, feeding results back, checking for more tool calls) until the model returns a final answer with no further tool requests, not stop after one exchange.
None of these providers is going to converge on identical syntax; each is optimizing tool calling around its own model's training and its own hosted-tool ecosystem. But the shape of a good integration, schema-first design, a normalization layer for the parts that differ, and a loop built for multi-step tool use rather than one-shot calls: is the same regardless of which model ends up on the other end of the request. Check the model comparison page if you're deciding which provider to build the integration around first.



