Skip to main content
Reasoning Prompts

Model Comparison Eval Generator

Given a task, generates a small test set you can run across multiple models to see which one actually performs best on your specific work.

intermediateWorks with any modelReasoning
Prompt
I need to compare how well different AI models handle a specific task. Generate a small evaluation set I can run with Promptfoo.

THE TASK: [DESCRIBE WHAT YOU WANT THE MODEL TO DO, e.g. "classify support tickets into billing/bug/feature_request/account" or "summarize a legal clause in plain English"]

MODELS I'M COMPARING: [e.g. "GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash" or "just testing effort levels on one model"]

WHAT "GOOD" LOOKS LIKE: [describe the quality bar — correctness, tone, format, whatever matters for this task]

Generate:
1. 8-10 test cases covering: 3-4 straightforward/typical inputs, 2-3 edge cases (ambiguous, missing information, unusually long or short), 1-2 adversarial cases (inputs designed to trip up a weak model)
2. For each test case, the expected output or an assertion that would catch a wrong answer (exact match, contains a specific phrase, or an llm-rubric description for subjective quality)
3. A ready-to-run `promptfooconfig.yaml` wired up with the providers I named and these test cases

How to use

Use this before committing to a model or a model tier for a specific task, rather than guessing based on general reputation. The output plugs directly into the site's Promptfoo tutorial workflow: save the generated config, run promptfoo eval, and you have a side-by-side comparison in minutes instead of an afternoon of manual testing.

Variables

  • [DESCRIBE WHAT YOU WANT THE MODEL TO DO] — be specific; "summarize" behaves very differently depending on what counts as a good summary
  • [MODELS I'M COMPARING] — name exact models (see the models page for current IDs), or specify you're comparing effort/thinking levels on a single model instead
  • [WHAT "GOOD" LOOKS LIKE] — the more concrete this is, the better the generated assertions will be

Tips

  • 8-10 cases is a starting point, not a final eval. If the comparison result surprises you (a model you expected to win doesn't), that's a signal to expand the test set for that specific task before trusting the result.
  • Adversarial cases matter more than they seem. A model that handles typical inputs well but falls over on a genuinely ambiguous one is a real production risk if your actual traffic includes ambiguous inputs, which it usually does.
  • Re-run this whenever you're about to move a task to a cheaper model tier. The model routing post covers the broader strategy; this prompt gets you the specific eval to justify (or block) a downgrade for your task.