Skip to main content
Model Guides

AI Model Guides

Each AI model has distinct strengths, quirks, and optimal prompting patterns. These guides cover what makes each model unique and how to get the best results from it.

Current models

Specs from each vendor's docs · verified

ModelContextMax outputPrice / 1M tokens (in / out)Best for
GPT-6 AstraOpenAI · gpt-6-astra1.05M128K$10 / $50Hardest reasoning and long-horizon agentic work
GPT-6 SolOpenAI · gpt-6-sol1.05M128K$2 / $10Everyday production workloads
GPT-6 LunaOpenAI · gpt-6-luna1.05M128K$0.10 / $0.50High-volume classification and extraction
Claude Fable 5.1Anthropic · claude-fable-5-11M128K$10 / $50Anthropic's most capable model for demanding long-running work
Claude Opus 5.5Anthropic · claude-opus-5-51M128K$4 / $20Coding and agentic work at a lower price than Fable
Claude Opus 5Anthropic · claude-opus-51M128K$5 / $25General-purpose Opus for reasoning, coding and agents
Claude Sonnet 5Anthropic · claude-sonnet-51M128K$2 / $10Daily driver for most production tasks
Claude Haiku 4.5Anthropic · claude-haiku-4-5200K—$1 / $5Fast, cheap sub-agents and high-volume routes
Gemini 3.8 FlashGoogle · gemini-3.8-flash1.05M65K$0.75 / $3.75Promo through Dec 31, 2026; $1.50 / $7.50 from Jan 1, 2027Fast multimodal work, including native video input
Gemini 3.5 Flash-LiteGoogle · gemini-3.5-flash-lite1.05M65K$0.30 / $2.50Cheapest multimodal option for high-volume routes

Prompting guides

How to Prompt Claude: Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5

Claude responds uniquely well to XML tags, explained constraints and clear output specs. The prompting patterns, thinking and effort settings, and model choices that work best with Claude in 2026.

ClaudeAnthropicXML tags
8 min read

How to Prompt GPT-6: Astra, Sol and Luna With the Responses API

OpenAI's GPT-6 family (Astra, Sol and Luna) shares a 1.05M-token context window. How to pick a tier, set reasoning effort, get structured outputs, and call tools with the Responses API.

OpenAIGPT-6Responses API
7 min read

How to Prompt Gemini 3.8 Flash: Video, Long Context, Thinking and Grounding

Gemini 3.8 Flash and 3.5 Flash-Lite take text, images, video, audio and PDFs in a 1M-token window. How to prompt them, set thinking levels, and use Search grounding and code execution.

GeminiGooglelong context
7 min read

How to Prompt LLaMA 3: Local Inference and Ollama SetupLegacy

How to run Meta's LLaMA 3 locally with Ollama, write effective prompts for it, and decide when a local model beats a hosted API. Kept for reference; newer open-weight models have since overtaken it.

LLaMAMetalocal LLM
5 min read

How to Prompt Mistral: Instruct Format, Efficiency, and API TipsLegacy

Mistral's model family balances strong performance with exceptional efficiency. Learn the instruct format, how to use the Mistral API, and when each model in the family fits best.

Mistralopen sourceefficiency
5 min read

GPT-6 vs Claude vs Gemini: Which AI Model for Which Task? (2026)

A practical comparison of GPT-6 Astra, Sol and Luna, Claude Fable 5.1, Opus 5.5 and Sonnet 5, and Gemini 3.8 Flash: context, inputs, reasoning controls and cost for real workloads.

model comparisonOpenAIClaude
5 min read

Not sure which model to use?

The comparison guide breaks down OpenAI, Claude, Gemini and open-weight models side by side.

See the comparison