OpenAI Evals is going away. The OpenAI Evals shutdown happens in two steps: on October 31, 2026 every existing eval becomes read-only, and on November 30, 2026 the Evals dashboard and API are scheduled to shut down entirely. The deprecation was announced back on June 3, so if you're reading this in October, the clock is already loud.
The dates come straight from OpenAI's deprecations page. What that page doesn't give you is an export button. The official migration guide says plainly that it "does not require an OpenAI Evals export feature" and then tells you to recreate your evals in Promptfoo.
That's fine if you have three evals. It's not fine if you've spent a year building graders, collecting run history and using those pass rates to decide which model ships. This guide covers the part the official docs skip: pulling everything out through the API while you still can, then rebuilding it in a form you own.
What the OpenAI Evals shutdown actually removes
Everything on the platform side goes:
- The Evals dashboard: eval definitions, run history, per-item results and report pages
- The Evals API: every endpoint under
/v1/evals - Eval graders: the testing criteria you configured (
string_check,text_similarity,score_model,label_model,python) - Dataset workflows that were part of the Evals product
What doesn't go: your OpenAI models, the Responses API, and anything you've logged in your own systems. The fine-tuning timelines are handled separately on the same deprecations page, so check those if you use graders for reinforcement fine-tuning.
The important date isn't November 30. It's October 31. After that you can still read your evals, but you can't create runs, so any comparison you want to rerun on the old platform has to happen before then.
Step 1: Export everything before October 31
Your data is reachable through the same API the dashboard uses. The Python SDK exposes three list calls that paginate automatically when you iterate over them:
client.evals.list(): eval definitions, includingtesting_criteriaanddata_source_configclient.evals.runs.list(eval_id): every run of an eval, with model, status andresult_countsclient.evals.runs.output_items.list(run_id, eval_id=...): every graded row, with the input (datasource_item), the model output (sample) and each grader's result
This script dumps all three into two files: one with definitions, one with every graded item.
import json
from openai import OpenAI
client = OpenAI()
with open("evals.jsonl", "w") as defs, open("eval-results.jsonl", "w") as results:
for ev in client.evals.list():
defs.write(json.dumps(ev.model_dump()) + "\n")
for run in client.evals.runs.list(ev.id):
for item in client.evals.runs.output_items.list(run.id, eval_id=ev.id):
results.write(json.dumps({
"eval_id": ev.id,
"eval_name": ev.name,
"run_id": run.id,
"run_name": run.name,
"model": run.model,
"datasource_item": item.datasource_item,
"sample": item.sample.model_dump(),
"results": [r.model_dump() for r in item.results],
"status": item.status,
}) + "\n")
print("Done. Keep both files in version control.")
Run it once now and again right before October 31 if you're still adding runs. Two things to know:
- It can take a while. Every output item is a separate row. An eval with 500 test cases and 40 runs is 20,000 rows. Leave it running.
datasource_itemis the gold. It's the exact input row each test used, including any reference answer you stored. That's what becomes your new test set. Thesamplefield holds the model's historical output, which is useful as a baseline but not something you'll re-grade.
If you only ever used the dashboard and uploaded datasets as files, check the Files section of the platform too. Those files aren't part of the Evals API responses, so download any you still need.
Step 2: Understand how the concepts map
Promptfoo and OpenAI Evals are built from the same pieces: test data, a prompt, a model, and rules for scoring the output. The difference is where they live. Evals kept everything in the OpenAI dashboard. Promptfoo keeps it in a promptfooconfig.yaml file in your repo, run from the CLI or in CI.
| OpenAI Evals | Promptfoo |
|---|---|
| Eval (dashboard object) | promptfooconfig.yaml in your repo |
| Data source / dataset rows | tests (inline, or file://tests.jsonl) |
datasource_item fields | vars, used in prompts as {{field}} |
| Model in a run | providers, e.g. openai:gpt-6-sol |
| Testing criteria / graders | assert entries on each test or in defaultTest |
| Run | promptfoo eval |
| Report page | promptfoo view |
And the graders, one by one:
| OpenAI grader | Promptfoo assertion | Notes |
|---|---|---|
string_check with eq | equals | Exact match |
string_check with like | contains | Substring match |
string_check with ilike | icontains | Case-insensitive substring |
text_similarity | similar | Embedding similarity with a threshold |
score_model | llm-rubric | Write the scoring rubric as the assertion value |
label_model | llm-rubric | Phrase the allowed labels as a rubric ("Output must be one of…") |
python | python | Port the grader function; the signature changes |
The model-graded ones deserve care. A score_model grader had its own prompt and pass threshold. When you rewrite it as an llm-rubric, run it on 20 or 30 rows where you already know the right answer and check that the new judge agrees with the old one. The official guide says the same thing: recreate each grader, then "verify its behavior before relying on it."
Step 3: Convert your exported rows into Promptfoo tests
Promptfoo reads tests from JSONL, one test per line, each with vars and assert. This turns the export into that format. It assumes your rows stored the reference answer in a field called expected. Change EXPECTED_FIELD to whatever you used.
import json
EVAL_NAME = "support-ticket-classifier"
EXPECTED_FIELD = "expected"
seen = set()
with open("eval-results.jsonl") as src, open("tests.jsonl", "w") as out:
for line in src:
row = json.loads(line)
if row["eval_name"] != EVAL_NAME:
continue
item = row["datasource_item"]
key = json.dumps(item, sort_keys=True)
if key in seen: # the same input appears once per run
continue
seen.add(key)
test = {"vars": item, "assert": []}
if EXPECTED_FIELD in item:
test["assert"].append({"type": "equals", "value": item[EXPECTED_FIELD]})
out.write(json.dumps(test) + "\n")
The de-duplication matters. Your export has one row per test per run, so without it an eval with 40 runs would give you 40 copies of every test.
Step 4: Rebuild the eval in Promptfoo
Here's a complete config for a ticket classifier that used a string_check grader plus a score_model grader for the explanation:
description: Support ticket classifier (migrated from OpenAI Evals)
prompts:
- |
Classify this support ticket as one of: billing, bug, feature_request, account.
Reply with the label on the first line, then one sentence explaining why.
Ticket: {{ticket}}
providers:
- openai:gpt-6-sol
- openai:gpt-6-luna
defaultTest:
options:
provider: openai:gpt-6-sol
assert:
- type: llm-rubric
value: The explanation refers to specific details from the ticket rather than generic wording.
tests: file://tests.jsonl
Two things changed on purpose. The prompt refers to {{ticket}}, which matches the field name in your old dataset rows. And there are two providers, so every run compares the model you use today with a cheaper one. That was clunky on the dashboard. Here it's one extra line.
Your label check came through as an equals assertion in tests.jsonl, but this prompt puts the label on the first line and an explanation after it. An exact match on the whole output will fail every row. Either change those assertions to a regex anchored at the start (^billing), or drop the explanation from the prompt and keep equals.
Then validate the config, run it and open the viewer:
promptfoo validate config -c promptfooconfig.yaml
promptfoo eval -c promptfooconfig.yaml --no-cache
promptfoo view
--no-cache forces fresh model calls, which is what you want the first time. After that you can drop it and let Promptfoo cache responses while you iterate on assertions.
If you haven't used Promptfoo before, our Promptfoo tutorial covers the rest: assertion types, CI with GitHub Actions and red-teaming.
Step 5: Check the new setup against your old numbers
Don't delete anything yet. You exported every historical run, and that's your check that the migration worked.
Take the last run you trust from the export and count its pass rate:
import json
passed = total = 0
with open("eval-results.jsonl") as f:
for line in f:
row = json.loads(line)
if row["run_name"] == "baseline-sept":
total += 1
passed += row["status"] == "pass"
print(f"Old platform: {passed}/{total} passed ({passed / total:.0%})")
Now run the same model in Promptfoo. The two numbers won't match exactly, since model-graded assertions and a re-run model both add noise, but they should be close. If you see a gap of more than a few points, it's almost always one of three things:
- A rubric that grades harder or softer than the old
score_modelprompt. Paste the old grader prompt into the rubric and compare row by row. - A missing variable. A
{{field}}in the prompt that doesn't exist invarsrenders as empty, and the model answers a different question. - Whitespace and casing.
equalsis strict. If the old grader usedilike, you wanticontains.
What you gain by moving
The shutdown is annoying, but it pushes evals where they should have been all along: next to the code.
When your eval config lives in the repo, a pull request that changes a prompt can also show the test results for that change. The review becomes "this prompt drops the ticket classifier from 94% to 89%" instead of "looks good to me." That's the workflow in our post on prompt versioning in production.
You're also no longer tied to one vendor's grading UI. The same tests.jsonl can score GPT-6 Sol against Claude Sonnet 5 or Gemini 3.8 Flash by adding providers. Current prices for each are on our models page.
And the thing that made your evals valuable was never the dashboard. It was the dataset. If yours is thin, the golden test set guide covers how many examples you need and how to label them.
Your checklist
- Run the export script and commit
evals.jsonlandeval-results.jsonl - Download any dataset files you uploaded to the platform
- Re-run any final comparisons you need on the old platform before October 31
- Convert each eval you still use to a
promptfooconfig.yaml - Check every model-graded rubric against 20–30 known rows
- Compare pass rates with your last trusted historical run
- Add
promptfoo evalto CI so it runs on prompt changes - Remove any code that calls
/v1/evalsbefore November 30



