Reasoning Prompts
Model Comparison Eval Generator
Given a task, generates a small test set you can run across multiple models to see which one actually performs best on your specific work.
Prompt
I need to compare how well different AI models handle a specific task. Generate a small evaluation set I can run with Promptfoo. THE TASK: [DESCRIBE WHAT YOU WANT THE MODEL TO DO, e.g. "classify support tickets into billing/bug/feature_request/account" or "summarize a legal clause in plain English"] MODELS I'M COMPARING: [e.g. "GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash" or "just testing effort levels on one model"] WHAT "GOOD" LOOKS LIKE: [describe the quality bar — correctness, tone, format, whatever matters for this task] Generate: 1. 8-10 test cases covering: 3-4 straightforward/typical inputs, 2-3 edge cases (ambiguous, missing information, unusually long or short), 1-2 adversarial cases (inputs designed to trip up a weak model) 2. For each test case, the expected output or an assertion that would catch a wrong answer (exact match, contains a specific phrase, or an llm-rubric description for subjective quality) 3. A ready-to-run `promptfooconfig.yaml` wired up with the providers I named and these test cases
How to use
Use this before committing to a model or a model tier for a specific task, rather than guessing based on general reputation. The output plugs directly into the site's Promptfoo tutorial workflow: save the generated config, run promptfoo eval, and you have a side-by-side comparison in minutes instead of an afternoon of manual testing.
Variables
[DESCRIBE WHAT YOU WANT THE MODEL TO DO]— be specific; "summarize" behaves very differently depending on what counts as a good summary[MODELS I'M COMPARING]— name exact models (see the models page for current IDs), or specify you're comparing effort/thinking levels on a single model instead[WHAT "GOOD" LOOKS LIKE]— the more concrete this is, the better the generated assertions will be
Tips
- 8-10 cases is a starting point, not a final eval. If the comparison result surprises you (a model you expected to win doesn't), that's a signal to expand the test set for that specific task before trusting the result.
- Adversarial cases matter more than they seem. A model that handles typical inputs well but falls over on a genuinely ambiguous one is a real production risk if your actual traffic includes ambiguous inputs, which it usually does.
- Re-run this whenever you're about to move a task to a cheaper model tier. The model routing post covers the broader strategy; this prompt gets you the specific eval to justify (or block) a downgrade for your task.