AI
OpenAI Evals Shutdown: Migrate Your Evals to Promptfoo
October 202611 min read

OpenAI announced the deprecation of the Evals platform on 3 June 2026. Existing evals become read-only on 31 October 2026, and the Evals dashboard and API are scheduled to shut down on 30 November 2026. OpenAI recommends Promptfoo as the replacement.
Export three things per eval: the eval object with its data_source_config and testing_criteria, the latest completed run with its model and input_messages template, and that run's output items, which hold each original row as datasource_item along with the per-grader scores. Together they are enough to rebuild the suite and to check that the rebuilt graders agree with the old scores.
Only as history. promptfoo import accepts an OpenAI Evals dashboard export such as eval_items_*.jsonl and stores it as a historical eval record with the grader scores, but it does not support the API output-items list response and does not create a runnable config. OpenAI's cookbook describes recreating the evaluation manually in a promptfooconfig.yaml.
string_check eq, ne, like and ilike map to equals, not-equals, contains and icontains and give identical results. text_similarity cosine maps to similar, the other similarity metrics to bleu, gleu, meteor and rouge-n, and score_model or label_model graders to llm-rubric with a threshold. Python graders become python assertions with a get_assert(output, context) signature instead of grade(sample, item).
OpenAI announced the acquisition of Promptfoo on 9 March 2026 and committed to keep maintaining its open-source tools, while integrating the technology into its OpenAI Frontier enterprise platform. Keeping your eval configuration, test data and graders as plain files in git is the safest hedge either way, because they stay portable whatever happens to the hosted products.

Key Takeaway
OpenAI Evals becomes read-only on 31 October 2026, and its dashboard and API shut down on 30 November 2026. Export each eval's definition, latest run and dataset rows through the API, import the dashboard JSONL into promptfoo only as history, then rebuild every grader as a promptfoo assertion in a YAML suite that GitHub Actions runs.
The OpenAI Evals shutdown has two dates, and the first one is a month away. On 31 October 2026 existing evals become read-only; on 30 November 2026 the Evals dashboard and API are scheduled to shut down. A regression suite that guards a production prompt, for example an ERP helpdesk bot that classifies tickets and drafts billing replies, stops existing on that second date unless someone moves it. OpenAI's own deprecations page names promptfoo as the way forward.
This is a deadline-driven migration guide. It covers what to export before the read-only date and through which API calls, why promptfoo's import command gives you history but not a runnable suite, how each OpenAI grader type maps to a promptfoo assertion and where the scores will not match, and how to run the rebuilt suite as a merge gate in GitHub Actions. Every behaviour cited comes from OpenAI's documentation and cookbook or promptfoo's command-line reference.
OpenAI announced the deprecation of the Evals platform on 3 June 2026. The deprecations page lists three dates: the announcement, 31 October 2026 when existing evals become read-only, and 30 November 2026 when the Evals dashboard and API are scheduled to shut down. The graders documented for eval workflows are part of the same transition, so the grader definitions inside your testing_criteria are on the clock too, not only the dashboard around them.
The same page schedules the v1/prompts API and reusable prompt objects for shutdown on the same 30 November date. A team that kept both its prompts and its evals in the dashboard has two migrations with one deadline, and the cheapest order is to move the prompt text into the repository first, because the rebuilt eval needs to point at it.
Treat 31 October as the real deadline, not 30 November. OpenAI does not list which operations a read-only eval still permits, so plan as if you cannot edit an eval, add a grader or start a fresh baseline run after that date. Any export that depends on creating a new run has to happen this month.
An eval on the platform is spread across three objects, and exporting only the first is the most common mistake. The eval object holds data_source_config and testing_criteria. A run holds the model and the input_messages template that every row was rendered into. The output items of a run hold each original row as datasource_item, plus per-grader results with a score and a pass flag. The script below walks all three for every eval in a project and writes them to disk.
# export_openai_evals.py — run it while the API still answers.
# Evals are listed per project, so run it once per project key.
import json, pathlib
from openai import OpenAI
client = OpenAI()
OUT = pathlib.Path("evals-export")
for ev in client.evals.list(limit=100): # the SDK pages through every eval
folder = OUT / ev.id
folder.mkdir(parents=True, exist_ok=True)
# The definition: data_source_config + testing_criteria (the graders).
(folder / "eval.json").write_text(ev.model_dump_json(indent=2))
runs = [r for r in client.evals.runs.list(ev.id) if r.status == "completed"]
if not runs:
continue # an eval that never ran has no dataset to recover
latest = max(runs, key=lambda r: r.created_at)
# What eval.json does not hold: data_source.model and the
# input_messages template the rows were rendered into.
(folder / "run.json").write_text(latest.model_dump_json(indent=2))
with (folder / "tests.jsonl").open("w", encoding="utf-8") as tests, \
(folder / "baseline.jsonl").open("w", encoding="utf-8") as base:
for out in client.evals.runs.output_items.list(latest.id, eval_id=ev.id):
# datasource_item is the original row. Nesting it under vars.item
# lets the old "item." template references work unchanged.
tests.write(json.dumps({
"description": f"row {out.datasource_item_id}",
"vars": {"item": out.datasource_item},
}) + "\n")
# The old scores, per grader: your parity baseline later.
base.write(json.dumps({
"row": out.datasource_item_id,
"scores": {r.name: r.score for r in out.results},
}) + "\n")The script writes tests.jsonl directly in promptfoo's test format, one test per line with the original row nested under vars.item. That nesting is deliberate: OpenAI templates reference the item namespace, and keeping the same shape means the old prompt template and the old grader references carry over without a rename. The baseline file records the old score for every grader on every row, which is the only way to prove later that the port behaves like the original.
Do not plan around a one-click export that produces a working promptfoo config. The current version of OpenAI's cookbook guide describes recreating an evaluation manually and says the process does not require an OpenAI Evals export feature. What does exist is a history path: promptfoo import accepts an OpenAI Evals dashboard export named like eval_items_*.jsonl and stores it as a historical promptfoo eval record.
promptfoo's reference describes what that import preserves: each source item under vars.item, the grader values, pass and fail states where the export has them, and the model output and token usage when the export includes sample data. Grader rows with a score but no pass state stay score-only rather than being turned into failures. It also states that the import supports the dashboard JSONL export, not the API output-items list response.
# History: the dashboard JSONL export of a completed run.
npx promptfoo import eval_items_OutputDataItemStatusParam.ALL.jsonl
npx promptfoo view # imported runs appear as historical evals
# Wrong: saving GET /v1/evals/{eval_id}/runs/{run_id}/output_items to a
# file and importing that. promptfoo documents the dashboard JSONL only;
# the API list response is not a supported import shape.
# Right: the API export feeds the REBUILD (tests.jsonl, baseline.jsonl),
# the dashboard export feeds the ARCHIVE. Keep both.So keep both artefacts, for different reasons. The dashboard export is the archive: past runs you can still open in promptfoo view after the platform is gone, useful when someone asks how the bot scored before last quarter's model change. The API export from the previous section is the raw material for the suite you will actually run from now on.
OpenAI's graders guide defines four grader families used in evals: string check, text similarity, score model and Python code execution, plus a label model grader in the API reference. Each has a promptfoo assertion that does the same job, but only the string checks are exact equivalents. The third column is the one to read before trusting a ported score.
| OpenAI grader | promptfoo assertion | Parity with the old score |
|---|---|---|
| string_check, eq | equals | Exact. Both are case-sensitive and return pass or fail. |
| string_check, ne | not-equals | Exact. Any promptfoo assertion can be inverted with the not- prefix. |
| string_check, like and ilike | contains and icontains | Exact. like is a case-sensitive substring test, ilike the case-insensitive one. |
| text_similarity, cosine | similar | Close. Both embed with text-embedding-3-large by default; carry pass_threshold over as threshold. |
| text_similarity, bleu, gleu, meteor, rouge | bleu, gleu, meteor, rouge-n | Same metric, different implementation. Re-baseline before gating on it. |
| text_similarity, fuzzy_match | python assertion calling rapidfuzz | Exact if you call the same rapidfuzz function. levenshtein is not an equivalent. |
| score_model and label_model | llm-rubric with threshold | Not exact. A different judge prompt and model will score differently. |
| python, grade(sample, item) | python, get_assert(output, context) | Exact once the signature is rewritten; the grading logic ports as is. |
The rows split into two groups. Deterministic graders, the string checks and a python grader that keeps its logic, should reproduce the old score on every row, so any mismatch there is a porting bug. Model-graded rows will drift, and OpenAI's cookbook says so directly: similarity-based scoring may not produce identical results across systems, and a recreated LLM-as-a-judge grader should be validated before it drives regression decisions.
Take a concrete case: an ERP helpdesk eval with three criteria, a string check that the reply mentions a refund, a score model grader that judges whether the reply stays grounded in the policy text, and a python grader that compares the reply with a reference answer. The config below rebuilds it, and each line points back to the exported field it came from.
# evals/support-reply/promptfooconfig.yaml
# Ported from the dashboard eval "Support answer quality".
description: ERP helpdesk reply quality
prompts:
- file://prompt.json # run.json -> input_messages template
providers:
- id: openai:responses:gpt-6-astra # run.json -> data_source.model
tests: file://tests.jsonl # one test per old datasource_item
defaultTest:
options:
# Pin the judge. Unset, promptfoo picks one from the API keys present,
# and a judge you did not choose is a score you cannot compare.
provider: openai:gpt-5-mini
assert:
# string_check, operation ilike, reference "refund" -> icontains
- type: icontains
value: refund
metric: mentions_refund
# score_model, range [0, 1], pass_threshold 0.7 -> llm-rubric
- type: llm-rubric
value: >-
The reply answers the question using only the policy text in the
ticket, quotes the invoice number exactly, and does not promise a
refund date.
threshold: 0.7
metric: policy_grounded
# python grader -> python assertion (see graders/reference_similarity.py)
- type: python
value: file://graders/reference_similarity.py
metric: reference_similarity
# evals/support-reply/prompt.json — the old template, copied verbatim.
# [
# { "role": "system", "content": "You answer ERP billing tickets. ..." },
# { "role": "user", "content": "{{ item.ticket_text }}" }
# ]Two details matter more than they look. The judge provider is pinned at defaultTest level, because promptfoo's llm-rubric otherwise picks a grading model based on which API keys are present, and an unpinned judge makes this week's score incomparable with next month's. And when an llm-rubric carries a threshold, promptfoo enforces both the judge's own pass verdict and the threshold, so a row can fail with a decent score if the judge said no.
# Before: OpenAI python grader (and text_similarity fuzzy_match, which
# uses rapidfuzz too). Signature fixed by the platform:
# def grade(sample, item) -> float
# sample["output_text"], item["reference_answer"]
# After: graders/reference_similarity.py, a promptfoo python assertion.
from rapidfuzz import fuzz, utils
PASS_THRESHOLD = 0.8 # copy the old grader's pass_threshold, do not re-pick it
def get_assert(output: str, context) -> dict:
item = context["vars"]["item"] # the old "item" namespace, intact
# The same WRatio call, so each row's score is comparable to baseline.jsonl.
score = fuzz.WRatio(
output, item["reference_answer"], processor=utils.default_process
) / 100.0
return {
"pass": score >= PASS_THRESHOLD,
"score": score,
"reason": f"WRatio {score:.2f} against the reference answer",
}
# Wrong: porting fuzzy_match as
# - type: levenshtein
# threshold: 0.9
# levenshtein's threshold is a maximum edit distance, not a 0-1 similarity.
# Every long answer fails, and the suite looks like a model regression.The python grader is the easiest port and the easiest to get subtly wrong. The function body survives; the signature changes from grade(sample, item) returning a float to get_assert(output, context), which may return a boolean, a score or a result object with pass, score and reason. Returning the object is worth the extra lines, because the reason string is what a reviewer reads when the CI job fails.
Once the suite lives in the repository, it should run where prompt changes are reviewed. promptfoo eval exits with code 100 when at least one test fails or the pass rate falls below PROMPTFOO_PASS_RATE_THRESHOLD, and with 1 for any other error, which is exactly the contract a CI job needs.
# .github/workflows/evals.yml
name: evals
on:
pull_request:
paths: ["evals/**", "src/prompts/**"]
jobs:
promptfoo:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: "24"
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: npm ci # promptfoo pinned in devDependencies
- run: pip install rapidfuzz # for the ported python graders
- uses: actions/cache@v4
with:
path: ~/.cache/promptfoo
key: ${{ runner.os }}-promptfoo-${{ hashFiles('evals/**') }}
- name: Run ported evals
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
# Default is 100: a single failed row fails the job. Start from the
# pass rate the old dashboard eval actually had, then ratchet up.
PROMPTFOO_PASS_RATE_THRESHOLD: "95"
run: npx promptfoo eval -c evals/support-reply/promptfooconfig.yaml -o results.json
- if: always() # keep the evidence when the gate fails
uses: actions/upload-artifact@v4
with:
name: promptfoo-results
path: results.jsonThe threshold is a percentage and defaults to 100, meaning a single failed row fails the job. A suite ported from a dashboard eval that historically passed 96 percent of rows will block every pull request on day one under that default. Set the threshold from the baseline the export gave you, and raise it as the flaky rows are fixed or removed, rather than starting at zero tolerance and teaching the team to ignore a red check.
Pin promptfoo as a devDependency and call it through npm ci rather than npx promptfoo@latest. Graders are code; a release that changes how a metric is computed moves your scores without any change to your prompt, and that is the one regression a regression suite cannot explain to you.
The order below front-loads everything that depends on the platform still accepting work, and leaves the purely local steps for later.
Do not delete or rewrite a test because the ported grader disagrees with the old score. A disagreement is either a porting bug or a real difference between judges, and both deserve a written note in the pull request. Silently dropping rows is how a 200-row suite becomes a 140-row suite that passes.
OpenAI announced its acquisition of Promptfoo on 9 March 2026, committed to keep its open-source tools maintained, and said the technology would be integrated into OpenAI Frontier, its enterprise agent platform. The recommended replacement for a deprecated OpenAI product is therefore a tool now owned by the same vendor. That is not a reason to avoid it, but it is a reason to depend on the part you control.
The part you control is the configuration. A promptfooconfig.yaml, a tests.jsonl and a folder of python graders are plain files: they are diffed in review, versioned with the prompt they test, and they keep working with a different provider line if the model behind the bot changes. The suite that disappears in November is the one stored only in a hosted dashboard, and the safest outcome of this migration is that it is the last one that can disappear that way.
The rule to carry forward: an eval is test code, so it lives next to the code it tests. Export the definitions, runs and rows before 31 October, archive the dashboard history with promptfoo import, rebuild graders as assertions with deterministic ones matching exactly, and let a CI exit code, not a dashboard, decide whether a prompt change ships.
Sources