Skip to main content
All Model Guides
Model GuideOpenAIGPT-6Responses APIstructured outputsfunction calling

How to Prompt GPT-6: Astra, Sol and Luna With the Responses API

OpenAI's GPT-6 family (Astra, Sol and Luna) shares a 1.05M-token context window. How to pick a tier, set reasoning effort, get structured outputs, and call tools with the Responses API.

7 min readVerified

OpenAI's current lineup is one family in three sizes. GPT-6 Astra, Sol and Luna share the same context window, the same inputs and the same API. What you're choosing between is capability against price, a 100× spread from top to bottom.

This guide covers how to prompt the GPT-6 family through the Responses API: picking a tier, setting reasoning effort, getting reliable structured output, and calling tools.


The GPT-6 family

GPT-6 AstraGPT-6 SolGPT-6 Luna
Model IDgpt-6-astragpt-6-solgpt-6-luna
Context window1.05M tokens1.05M tokens1.05M tokens
Max output128K tokens128K tokens128K tokens
InputText, imageText, imageText, image
Price per 1M tokens (in / out)$10 / $50$2 / $10$0.10 / $0.50
Use it forHardest reasoning, long-horizon agentsEveryday production workClassification, extraction, routing at volume

OpenAI's reasoning guide recommends starting with Astra for most reasoning workloads. In practice, the cost-sensible path is the reverse: start with the cheapest tier that could plausibly work, test it, and move up only where it fails. Luna costs 1% of Astra per token.

Two things the family doesn't do: take audio or video as chat input (OpenAI has separate realtime and transcription models for audio), or output images (that's the GPT Image models). For native video understanding, see Gemini.


The Responses API

OpenAI's primary API for current models is Responses. Chat Completions still works, but OpenAI says reasoning models perform better on Responses.

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6-sol",
    instructions="You are a senior Python reviewer. Be direct and specific.",
    input="Review this function for bugs:\n\n[code]",
)
print(response.output_text)

instructions plays the role of the system prompt. input can be a plain string, or a list of messages and content parts when you need images or multi-turn history.


Reasoning effort

GPT-6 models reason internally before answering. You control how much with reasoning.effort:

response = client.responses.create(
    model="gpt-6-astra",
    reasoning={"effort": "high", "summary": "auto"},
    input="Design a zero-downtime migration plan for splitting our users table into two services. [context]",
)
print(response.output_text)
  • Effort values run from none through minimal, low, medium, high, xhigh and max. Which ones a model supports varies.
  • Reasoning tokens are hidden but billed as output tokens and take up context window.
  • summary (auto, concise or detailed) returns a readable summary of the reasoning in the response's reasoning item.

The prompting consequence matters most. OpenAI's guidance is to give the model "the task, constraints, and desired output format" and avoid prescribing intermediate steps. Drop "think step by step" and numbered reasoning scripts; raise effort instead. More in how to prompt reasoning models.


Prompting patterns that work

Describe the result, not the route.

Bad:  First read the ticket. Then decide the category. Then write a reply.

Good: Classify this support ticket as billing, bug, feature_request or account,
      and draft a reply under 120 words in our support voice.
      If the ticket mentions a refund over $500, set escalate to true.

Put hard constraints in instructions, task details in input. Instructions are for rules that apply to every request: tone, audience, what never to do. Keep them short and stable; that also helps prompt caching.

Say what "done" looks like. Length limits, required sections, what to do when information is missing ("reply unknown rather than guessing").

Ask for a check, not a transcript. "Verify the totals match the line items before answering" improves accuracy without paying for visible step-by-step output.


Structured outputs

When code will parse the answer, don't rely on "reply in JSON." Pass a schema and let the API enforce it. With the Python SDK, a Pydantic model is the easiest route:

from pydantic import BaseModel
from openai import OpenAI

client = OpenAI()

class Ticket(BaseModel):
    category: str
    priority: int
    escalate: bool
    summary: str

response = client.responses.parse(
    model="gpt-6-luna",
    input="Classify this ticket: 'I was charged twice for my annual plan and need a refund today.'",
    text_format=Ticket,
)

ticket = response.output_parsed
print(ticket.category, ticket.escalate)

The model's output is constrained to the schema, so you won't get a missing field or a stray sentence before the JSON. Classification like this is a good Luna job. Test it on your real tickets before assuming you need Sol.

Two tips that still apply from the GPT-4o era:

  • Enums beat free text. If category can only be four values, make it a Literal[...] so the model can't invent a fifth.
  • Name fields for meaning. escalate_to_human gets filled more accurately than flag2.

Function calling

Define tools as JSON Schema with "strict": True, and the model returns structured calls you execute:

import json
from openai import OpenAI

client = OpenAI()

tools = [{
    "type": "function",
    "name": "get_order_status",
    "description": "Look up the shipping status of an order by its ID.",
    "parameters": {
        "type": "object",
        "properties": {"order_id": {"type": "string"}},
        "required": ["order_id"],
        "additionalProperties": False,
    },
    "strict": True,
}]

input_items = [{"role": "user", "content": "Where is order A-1042?"}]
response = client.responses.create(model="gpt-6-sol", tools=tools, input=input_items)

# Send the model's output back, plus a result for each function call
input_items += response.output
for item in response.output:
    if item.type == "function_call":
        args = json.loads(item.arguments)
        result = get_order_status(**args)
        input_items.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(result),
        })

final = client.responses.create(model="gpt-6-sol", tools=tools, input=input_items)
print(final.output_text)

What makes tool use reliable:

  • Descriptions are prompts. "Look up the shipping status of an order by its ID" tells the model when to call the tool. Spend your effort there.
  • Handle several calls per turn. The model can request more than one function at once. Return a function_call_output for every call_id.
  • Parse arguments as JSON. Never string-match on the raw arguments text.

The concepts carry across vendors; see the function calling lesson.


Images

GPT-6 models take images alongside text:

response = client.responses.create(
    model="gpt-6-sol",
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": (
                "This is a screenshot of our checkout page. List every UI element "
                "a first-time user might find confusing, with a specific fix for each."
            )},
            {"type": "input_image", "image_url": "https://example.com/checkout.png", "detail": "auto"},
        ],
    }],
)

As with every vision model, the prompt decides the quality. "Describe this image" gets a caption. Name exactly what to find and what form the answer takes.


Common mistakes with GPT-6

Defaulting to Astra. It's the best model and 100× Luna's price. Run your task on Luna and Sol first. For extraction and classification the smaller tiers often pass.

Scripting the reasoning. Step-by-step instructions written for GPT-4-era models can make reasoning models worse. Specify the result and set effort.

Staying on Chat Completions for reasoning work. It still works, but OpenAI recommends Responses for reasoning models.

Prompt-only JSON. Use structured outputs (text_format or a JSON schema) whenever code consumes the result.

Ignoring deprecations. OpenAI retires older snapshots on fixed dates. Several legacy GPT snapshots shut down on October 23, 2026. Check the deprecations page for anything you pin. If you used the Evals dashboard, it shuts down on November 30, 2026; here's how to migrate.

Current prices for every model, including Claude and Gemini, are on the models page.

Want to compare models side by side?

See how OpenAI, Claude, Gemini and open-weight models stack up for different use cases.

View model comparison →