Somebody tidies the support-triage prompt. They delete a line that looks redundant: the one telling the model refund requests are billing. The PR is two lines, it reads fine, and a reviewer approves it. Three days later refund tickets are landing in the wrong queue.
A prompt change is a behaviour change with no compiler to complain about it. Prompt regression tests in CI are the compiler: a fixed set of inputs with expected outcomes, run on every PR that touches the prompt, failing the build when the pass rate drops.
This post is the CI half. The Promptfoo complete guide covers assertion types and the UI, and the OpenAI Evals migration post covers moving old evals over. Here: the dataset, a threshold you can defend, handling flaky results, keeping the bill sane, and a workflow file. I ran the plumbing locally with Promptfoo 0.124.1 against a deliberately fake model, so you'll see real exit codes and pass rates, and I'll be clear about what that does and doesn't prove.
What do I need in the repo?
Five files and one workflow.
prompts/triage.txt
tests/tickets.csv
promptfooconfig.ci.yaml # real provider, runs in CI
promptfooconfig.yaml # mock provider, runs anywhere for free
mock_provider.js
.github/workflows/prompt-eval.yml
The prompt is plain text with a {{ticket}} variable:
You triage support tickets. Return only JSON: {"category": "billing|bug|how_to|account", "priority": "low|high"}.
Rules:
- Anything about charges, invoices, or refund requests is billing.
- Crashes, errors, and data that looks wrong are bug.
- Questions about using a feature are how_to.
- Login, password, and SSO problems are account.
- Mark priority high if the customer cannot work at all.
Ticket:
{{ticket}}
Where does the dataset come from?
From your failures, not your imagination. Each bug report that traces back to the prompt becomes a row, with the input that triggered it and the answer you wanted. Promptfoo can load tests from a CSV: column headers become variables, and __description names the test (per its test-cases docs).
__description,ticket,expected_category
charged twice,"I was charged twice for March, please fix the invoice",billing
refund ask,I want a refund for the annual plan,billing
export crash,The export button crashes the app every time,bug
wrong totals,Dashboard totals look wrong after the import,bug
howto filters,How do I save a filter on the reports page?,how_to
howto api,Where do I find my API key to use the integration?,how_to
sso locked,SSO login loops back to the sign-in page,account
password reset,I never got the password reset email,account
Eight rows is a demo. For a real prompt, aim for 30 to 50 covering each category, plus awkward inputs: empty text, a non-English ticket, a ticket that contains instructions addressed to the model. How to build and grow that set properly is in building evaluation datasets and golden test sets.
Which assertions should gate the build?
Prefer ones that can't flake. The config below puts a JSON check and an exact-category check in defaultTest, so they apply to every row. context.vars gives the JavaScript assertion access to the row's expected_category; I confirmed that works by running it.
description: Support ticket triage prompt, regression suite
prompts:
- file://prompts/triage.txt
providers:
- id: file://./mock_provider.js
label: mock-model
evaluateOptions:
maxConcurrency: 4
defaultTest:
assert:
# Hard gates: output must be JSON and must stay cheap and short.
- type: is-json
- type: javascript
value: "JSON.parse(output).category === context.vars.expected_category"
metric: category_correct
- type: latency
threshold: 5000
tests: file://tests/tickets.csv
Add an LLM-judged llm-rubric for things string checks can't express, but sparingly, for two reasons from Promptfoo's docs. First, with no threshold set, a grader response with a missing pass field counts as a pass, so a score of 0 can slip through; set a threshold. Second, if you don't choose a judge, Promptfoo picks one from whichever credentials it finds (with an Anthropic key it uses claude-sonnet-5, per the docs I read today). That is a cost decision you didn't make. Pin it with defaultTest.options.provider.
What does the real-provider config look like?
This is the one CI uses. I ran promptfoo validate config on it (it reported the configuration valid), but I did not run it against the API, because I had no key in the sandbox. Treat it as validated, not exercised.
description: Support ticket triage, CI suite (real provider)
prompts:
- file://prompts/triage.txt
providers:
- id: anthropic:messages:claude-sonnet-5-5
label: triage-model
config:
max_tokens: 100
evaluateOptions:
maxConcurrency: 4
defaultTest:
options:
# llm-rubric picks its own judge from your credentials; pin it so cost is predictable.
provider:
id: anthropic:messages:claude-sonnet-5-5
config:
max_tokens: 300
assert:
- type: is-json
- type: javascript
value: "JSON.parse(output).category === context.vars.expected_category"
metric: category_correct
- type: latency
threshold: 8000
- type: llm-rubric
value: "The priority is 'high' only if the ticket says the customer cannot work at all."
threshold: 0.8
metric: priority_sane
tests: file://tests/tickets.csv
The provider format anthropic:messages:<model> and claude-sonnet-5-5 both appear in Promptfoo's Anthropic page. Use a cheaper model for the judge if your provider offers one; I used the same ID for both so every identifier here is one I saw documented in both places.
What does a regression look like?
To check the gate works I need a prompt that can regress, and a real model won't do that on demand. So mock_provider.js is a stand-in "model": it reads the rendered prompt and only applies a rule if the rule's text is still in the prompt. Delete the billing line and billing tickets stop being classified.
// Stand-in for a real model, used only to exercise the promptfoo plumbing.
// It "reads" the rendered prompt: a rule only works if the prompt text still contains it.
// FLAKY=1 makes one ticket fail on every third call, to demonstrate --repeat.
let calls = 0;
module.exports = class MockProvider {
constructor(options) {
this.providerId = (options && options.id) || 'mock';
}
id() {
return this.providerId;
}
async callApi(prompt) {
calls += 1;
const [rules, ticketPart] = prompt.split('Ticket:');
const t = (ticketPart || '').toLowerCase();
const has = (s) => rules.toLowerCase().includes(s);
let category = 'how_to';
if (/charged|invoice/.test(t) && has('charges, invoices')) category = 'billing';
else if (/refund/.test(t) && has('refund requests')) category = 'billing';
else if (/crash|wrong/.test(t)) category = 'bug';
else if (/sso|password/.test(t)) category = 'account';
if (process.env.FLAKY && t.includes('api key') && calls % 3 === 0) category = 'account';
return {
output: JSON.stringify({ category, priority: 'low' }),
tokenUsage: { total: 30, prompt: 20, completion: 10 },
cost: 0,
};
}
};
The custom-provider shape (file:// id, a class with id() and callApi(prompt, context, options) returning output and tokenUsage) is from Promptfoo's custom provider docs. I ran three scenarios with --no-cache:
| Scenario | Result | Exit code |
|---|---|---|
| Original prompt | 8 of 8 passed (100%) | 0 |
| Billing line deleted from the prompt | 6 of 8 passed (75%); the two failures were the charged-twice and refund tickets | 100 |
Same broken prompt, PROMPTFOO_PASS_RATE_THRESHOLD=70 | 6 of 8 passed (75%) | 0 |
Two things to take from that table. The exit code is the gate: Promptfoo's CLI docs say 100 means a test failed or the pass rate fell below the threshold, and 1 means any other error. And the threshold changes the meaning of the gate entirely. At the default (100) any single failure blocks the merge. At 70, a prompt that broke a quarter of its tests sails through. Pick the number on purpose, and write down why.
One more finding. Some guides, including an older post on this site, tell you to add --ci or --fail-on-error to make evals fail CI. On 0.124.1 both are rejected with unknown option, and promptfoo eval --help lists neither. The exit code behaviour above is what to rely on. If a different version behaves differently, --help will tell you.
How do I deal with flaky tests?
Decide what a flaky test is first. Real models give different answers to the same input, so some tests will pass 9 times out of 10. Re-running the job until it goes green is the worst response, because it trains everyone to ignore red.
What to do instead:
- Prefer deterministic assertions. A category match can't be argued with. A rubric judged by a second model adds its own variance on top of the first.
- Pin what you can. Promptfoo's Anthropic page says temperature defaults to 0, but also that on models without sampling controls it omits the parameter entirely, so you can't assume determinism for every model.
- Repeat and gate on an aggregate.
--repeat Nruns each test N times (it's ineval --help). With three repeats, one bad sample out of 24 doesn't have to block a merge.
I tested the third with the mock set to fail the API-key ticket on every third call. Running with --repeat 3 gave 24 results, 23 passed and 1 failed (95.83%). At the default threshold the run exited 100. Setting PROMPTFOO_PASS_RATE_THRESHOLD=90 would let that through, which is the policy question again: how many wrong answers per hundred will you tolerate? My view: a high but not perfect bar (90 to 95) for a suite with repeats, 100 for a small deterministic suite.
Then track the tests that flip. A test that fails one run in five is either a bad test or a prompt that's fragile on that input. Both are worth knowing about. Promptfoo's caching docs say each repeat index gets its own cache namespace, so cached reruns are repeatable per repeat instead of collapsing three samples into one.
How do I keep the bill down?
Count the calls first. Each test costs one model call per repeat, plus one judge call per llm-rubric per repeat. With 40 tests, --repeat 3 and one rubric, that's 40 × 3 × 2 = 240 calls per run. Multiply by the number of PR pushes and it adds up, so the levers matter more than the model choice:
- Path filters. The workflow only triggers when prompts, tests or the eval config change. Unrelated PRs cost nothing.
- Cancel superseded runs. A
concurrencygroup withcancel-in-progress: truestops paying for a run the next push has already replaced. - Cache responses. Promptfoo's cache is on by default (disk, 14-day TTL per its docs). Persist it with
PROMPTFOO_CACHE_PATHand the Actions cache keyed on a hash of the prompt, tests and config. Unchanged inputs are then free; a changed prompt busts the key. - Pin a cheap judge. As above.
- Sample on PRs, run everything on a schedule.
--filter-sample Nwith--filter-sample-seedpicks a repeatable random subset (both flags are in--help; I didn't use them in the workflow below, and a nightly full run needs a second workflow with ascheduletrigger that I haven't written here). - Limit concurrency.
--max-concurrency(also-j) helps with provider rate limits as much as with cost.
Check actual spend on your provider's usage page after the first few runs. I'm not putting a dollar figure on a suite I haven't run against a real model.
The workflow
name: Prompt regression
on:
pull_request:
paths:
- "prompts/**"
- "tests/**"
- "promptfooconfig.ci.yaml"
- ".github/workflows/prompt-eval.yml"
permissions:
contents: read
concurrency:
group: prompt-eval-${{ github.event.pull_request.number }}
cancel-in-progress: true
env:
PROMPTFOO_CACHE_PATH: ${{ github.workspace }}/.promptfoo-cache
PROMPTFOO_DISABLE_TELEMETRY: "1"
jobs:
eval:
# Fork PRs get no secrets, so the real-provider run can't work for them.
if: github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-node@949feb2413d6458794dcd2491c4babbbce0c15c1 # v7.1.0
with:
node-version: "24"
- name: Restore response cache
id: cache
uses: actions/cache/restore@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: .promptfoo-cache
key: promptfoo-${{ hashFiles('prompts/**', 'tests/**', 'promptfooconfig.ci.yaml') }}
- name: Run evals
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
PROMPTFOO_PASS_RATE_THRESHOLD: "90"
run: >-
npx --yes promptfoo@0.124.1 eval
-c promptfooconfig.ci.yaml
--repeat 3
--max-concurrency 4
--no-progress-bar --no-table --no-share
-o results.json
# actions/cache only saves when the job succeeds; a failing eval is the run you most want cached.
- name: Save response cache
if: always() && steps.cache.outputs.cache-hit != 'true'
uses: actions/cache/save@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: .promptfoo-cache
key: promptfoo-${{ hashFiles('prompts/**', 'tests/**', 'promptfooconfig.ci.yaml') }}
- name: Summarize failures
if: always()
run: |
test -f results.json || exit 0
{
echo "### Prompt regression"
jq -r '.results.stats | "passed \(.successes), failed \(.failures), errors \(.errors)"' results.json
echo
jq -r '.results.results[] | select(.success == false) | "- \(.vars.ticket) (expected \(.vars.expected_category))"' results.json | sort | uniq -c
} >> "$GITHUB_STEP_SUMMARY"
Choices worth explaining:
- Pinned version.
promptfoo@0.124.1is the version I ran, so the flags and exit codes above are the ones in this post. Bump it deliberately, run once, and read--help. - Node 24. Promptfoo's GitHub Action page says its runner needs Node 22.22 or newer and recommends 24 LTS. I used 24 here and did not check what the CLI itself requires.
- No
--share.--no-sharekeeps results out of any shared link. Check what you're comfortable sending before turning sharing on. - Fork PRs are skipped. GitHub doesn't pass secrets to workflows from forks, so the job couldn't call the API anyway. I'd rather skip than reach for
pull_request_target, which GitHub's security guidance says to avoid unless necessary. The same reasoning is spelled out in the AI code review workflow post. - Summary step. The
jqfilters readresults.json. I ran the twojqfilters (without thesortanduniqat the end) on the regressed run's output and they printedpassed 6, failed 2, errors 0and the two billing tickets with their expected categories. I did not run the whole workflow on GitHub; I parsed it with PyYAML and checked action SHAs against the actions' own repositories today (re-check when you copy them). - Promptfoo's own action.
promptfoo/promptfoo-action@v1exists and posts PR comments, with inputsgithub-token,prompts,config,openai-api-keyandcache-pathper its docs. I used the CLI because its docs list an OpenAI key input only, and I wanted the exit code and threshold visible in the YAML. If you use OpenAI and want comments on the PR, look at it.
What does this not tell you?
Be honest about what green means. It means: on these inputs, the prompt and model produce output that passes these assertions. It does not mean the prompt is good. The run I show uses a fake model; it proves the gate, the threshold, the repeat handling and the reporting work, not that your prompt is safe to ship. Passing tests also only cover inputs you thought of, so keep feeding production failures back into the CSV.
Version your prompts like code so a failing run points at a diff: prompt versioning in production covers that. And if your system is an agent rather than a single prompt, you'll want task-level checks, which agent evals from scratch walks through.
Before you merge it
- Run the mock config locally and confirm you get exit 0, then break the prompt and confirm exit 100.
- Run
promptfoo validate config -c promptfooconfig.ci.yamland one real run with your key. Read the failures before choosing a threshold. - Write the threshold and the reason in the PR description.
- Check that the pinned judge is the one you meant to pay for.
- Open a test PR that edits only a README. The workflow should not run.



