OpenAI's billing page has a special talent for rejecting Indian cards. Domestic debit cards don't work at all. Credit cards work only if your bank has international online transactions explicitly enabled — and even then, some banks block it at the card network level. The workarounds people use: Wise card (requires opening a foreign currency account), a US-based friend's card (doesn't scale), or just giving up.
If you want to use GPT-6 via the API and you're in India, AICredits.in is the most straightforward path. It's an OpenAI-compatible gateway that routes your API calls and bills you in INR via Razorpay. UPI, net banking, domestic cards — all of them work.
Supported OpenAI models
OpenAI's current model IDs are gpt-6-astra, gpt-6-sol and gpt-6-luna. Estimated INR pricing below is calculated from OpenAI's published USD list prices plus AICredits' stated markup (5% forex buffer + 5% platform fee, explained further down) at a rate of roughly ₹95.7/USD — check the live AICredits dashboard for the exact current rate, since both the exchange rate and OpenAI's list prices change.
| Model | Model ID | Input (est. INR/1M tokens) | Output (est. INR/1M tokens) |
|---|---|---|---|
| GPT-6 Astra | openai/gpt-6-astra | ~₹1,053 | ~₹5,264 |
| GPT-6 Sol | openai/gpt-6-sol | ~₹211 | ~₹1,053 |
| GPT-6 Luna | openai/gpt-6-luna | ~₹11 | ~₹53 |
GPT-6 Luna at roughly ₹11/1M input tokens is worth calling out. For high-volume tasks — classification, extraction, summarization, routing — it's the right default. That's roughly ₹11 to process around 100 articles. That's not a rounding error, that's genuinely cheap.
Migration from direct OpenAI: 2 line changes
If you're already using the OpenAI SDK, migration is:
# Before
from openai import OpenAI
client = OpenAI(api_key="sk-your-openai-key")
# After
from openai import OpenAI
client = OpenAI(
api_key="sk-your-aicredits-key",
base_url="https://api.aicredits.in/v1"
)
That's it. Every call you make with this client now routes through AICredits. The model names get a provider prefix (openai/gpt-6-sol instead of gpt-6-sol), but that's the only other change.
Before:
response = client.chat.completions.create(
model="gpt-6-sol",
...
)
After:
response = client.chat.completions.create(
model="openai/gpt-6-sol", # prefix added
...
)
Chat completions
Full working example with system prompt, multi-turn conversation, and temperature:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AICREDITS_API_KEY"],
base_url="https://api.aicredits.in/v1"
)
def chat(messages: list, model: str = "openai/gpt-6-luna") -> str:
response = client.chat.completions.create(
model=model,
messages=messages,
temperature=0.7,
max_tokens=1024
)
return response.choices[0].message.content
# Multi-turn conversation
history = [
{"role": "system", "content": "You are a senior backend engineer. Be terse and specific."}
]
history.append({"role": "user", "content": "What's the best way to handle database connection pooling in Python?"})
reply = chat(history)
history.append({"role": "assistant", "content": reply})
print(reply)
history.append({"role": "user", "content": "How does this change if I'm using async SQLAlchemy?"})
reply = chat(history)
print(reply)
Streaming
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AICREDITS_API_KEY"],
base_url="https://api.aicredits.in/v1"
)
with client.chat.completions.create(
model="openai/gpt-6-sol",
messages=[
{"role": "user", "content": "Write a FastAPI endpoint that accepts a JSON body, validates it with Pydantic, and writes to PostgreSQL using asyncpg"}
],
stream=True,
max_tokens=2048
) as stream:
for chunk in stream:
delta = chunk.choices[0].delta
if delta.content:
print(delta.content, end="", flush=True)
print()
The streaming response format is identical to OpenAI's SSE format. Anything that consumes that format — Vercel AI SDK, LangChain streaming callbacks, your own SSE parser — works without modification.
Function calling (tool use)
Function calling works exactly as it does with direct OpenAI:
import json
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AICREDITS_API_KEY"],
base_url="https://api.aicredits.in/v1"
)
tools = [
{
"type": "function",
"function": {
"name": "get_stock_price",
"description": "Get the current stock price for an Indian company",
"parameters": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "NSE ticker symbol, e.g. RELIANCE, TCS, INFY"
}
},
"required": ["ticker"]
}
}
}
]
response = client.chat.completions.create(
model="openai/gpt-6-sol",
messages=[{"role": "user", "content": "What's the current price of TCS stock?"}],
tools=tools,
tool_choice="auto"
)
message = response.choices[0].message
if message.tool_calls:
tool_call = message.tool_calls[0]
function_name = tool_call.function.name
arguments = json.loads(tool_call.function.arguments)
print(f"Model wants to call: {function_name}({arguments})")
# → Model wants to call: get_stock_price({'ticker': 'TCS'})
This is the foundation of agentic prompting — letting the model decide when to call external tools rather than hardcoding the flow. Build this pattern right and you can swap GPT-6 for Claude or Gemini without touching the tool definitions.
Reasoning effort
Every GPT-6 tier reasons internally now — there's no separate reasoning-only model to reach for. You control depth with a reasoning_effort field, passed the same way through the gateway:
response = client.chat.completions.create(
model="openai/gpt-6-sol",
messages=[
{
"role": "user",
"content": """A startup has ₹50,000 monthly budget. They need:
- 10,000 GPT-6 Sol calls averaging 1,000 input tokens and 500 output tokens each
- 50,000 GPT-6 Luna calls averaging 500 input tokens and 200 output tokens each
Calculate total cost using AICredits pricing and whether they fit in budget."""
}
],
# reasoning_effort: "low" | "medium" | "high" | "xhigh" | "max"
extra_body={"reasoning_effort": "medium"}
)
print(response.choices[0].message.content)
For coding problems, SQL queries, and structured reasoning tasks, raising reasoning_effort on Sol is often better value than jumping straight to Astra. Drop "think step by step" from the prompt — the model already reasons before answering — and see how to prompt reasoning models for more.
Image generation
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AICREDITS_API_KEY"],
base_url="https://api.aicredits.in/v1"
)
response = client.images.generate(
model="openai/gpt-image-2.5-flare",
prompt="A Bangalore tech office at dusk, floor-to-ceiling glass windows, city lights below, warm interior lighting, cinematic photography style",
size="1024x1024",
quality="standard",
n=1
)
image_url = response.data[0].url
print(image_url)
Check the AICredits dashboard for the current image model catalog and INR pricing — image generation is billed per image rather than per token, and OpenAI's image lineup changes independently of the chat models.
Automatic failover
One underappreciated feature: if the underlying OpenAI endpoint is throttling or returning errors, AICredits can route to a backup. This isn't guaranteed instant recovery, but for production workloads running overnight batch jobs or handling user traffic, it reduces the blast radius of an OpenAI outage.
In practice this means your error rate stays lower than if you called OpenAI directly during a partial outage. For teams running customer-facing features on top of GPT-6, this matters.
Common gotcha: model name format
The single most common error when switching to AICredits:
Error: model "gpt-6-sol" not found
Fix: prefix with the provider.
# Wrong
model="gpt-6-sol"
# Right
model="openai/gpt-6-sol"
Set this in an environment variable or constant so you only have to remember it once:
# config.py
GPT6_ASTRA = "openai/gpt-6-astra"
GPT6_SOL = "openai/gpt-6-sol"
GPT6_LUNA = "openai/gpt-6-luna"
CLAUDE_HAIKU = "anthropic/claude-haiku-4-5"
GEMINI_FLASH = "google/gemini-3.8-flash"
Then use config.GPT6_SOL in your calls. When you want to experiment with a different model, you change one line.
Per-key budget caps
Create separate API keys for separate projects or team members. Set a budget cap on each key in the AICredits dashboard. When the cap is hit, that key stops working — no runaway costs from a bug in production.
Recommended structure for a small team:
- Production key: Higher cap (₹5,000/month), tight access control
- Dev key: ₹500 cap, shared with developers
- Experiment key: ₹200 cap, used for testing new models or prompts
The dashboard shows per-key usage breakdowns so you can see exactly what each project is spending.
What you're getting for the 10% markup
The AICredits fee is transparent: 5% forex buffer + 5% platform fee on top of live rates. Consider what you'd otherwise pay: a Wise card has exchange rate spreads of 0.5–1.7% plus a monthly fee; an international credit card charges 2–3.5% as a foreign transaction fee. By the time you factor in those costs plus the hassle of maintaining a foreign currency account, 10% is often cheaper — and the UPI/domestic card access is priceless if you don't have that infrastructure already.
For the actual implementation patterns that consume most of your API budget, prompting for coding covers the prompts that give the highest quality-to-token ratio on code tasks. If you're building agents on top of GPT-6, the Agents track covers the architecture decisions that matter. Current USD list prices for every model are on the models page.



