Your agent said "Order cancelled." The order is still open. If your eval reads the reply, that run is a pass.
This post builds the smallest agent eval harness I could make useful, in about 120 lines of standard-library Python: tasks with a known right outcome, a grader that checks what actually changed in the world, rules about which tool calls are allowed, and repeated trials so you see how often it works instead of whether it worked once. Then I run it and show the numbers, including a trap in the numbers.
One limit up front: the agent in this post is a stub, not an LLM. It's a Python function with seeded random failures I chose on purpose (a retry bug, a skipped tool call, obeying the user too literally). That is deliberate. It lets the harness be deterministic and runnable, and it lets me know the true failure rates so I can show you when the measurement is wrong. Swapping in a real model is a ten-line change, covered below, and I have not run that version.
If you want the wider picture of what to measure in production (latency, cost, monitoring, LLM-as-judge), the earlier post on AI agent evaluation frameworks covers it. For how to source and label input sets, see building evaluation datasets and golden test sets. This one is narrower: the mechanics of grading an agent that takes actions.
What makes agent evals different from prompt evals?
A prompt eval compares text to a reference. An agent eval has to answer three separate questions, and one score can't hold all of them:
- Did the world end up right? The refund exists. The order is cancelled. The file was written.
- Did it get there acceptably? No forbidden tools. No double refunds. No emails nobody asked for. Calls in a sane order.
- Does it do this reliably? Once is an anecdote.
Text similarity answers none of these. A transcript can read perfectly and the database be wrong, or the database can be right after the agent tried something it should never have attempted.
Step 1: build a tiny world with real state
The agent needs something to act on, and the grader needs something to inspect. Mine is an in-memory shop with three orders and four tools. Every call is recorded, so tool-use checks come free.
import copy
import random
from math import comb
INITIAL = {
"orders": {
"A100": {"status": "delivered", "total": 40, "refunded": 0},
"A101": {"status": "shipped", "total": 25, "refunded": 0},
"A102": {"status": "delivered", "total": 90, "refunded": 0},
},
"emails": [],
}
class World:
def __init__(self):
self.db = copy.deepcopy(INITIAL)
self.calls = [] # every tool call, in order
def call(self, tool, **args):
self.calls.append((tool, args))
o = self.db["orders"].get(args.get("order_id"))
if tool == "lookup_order":
return o or {"error": "not found"}
if tool == "issue_refund":
if not o or o["status"] != "delivered":
return {"error": "not refundable"}
o["refunded"] += args["amount"]
return {"ok": True}
if tool == "cancel_order":
if not o or o["status"] != "shipped":
return {"error": "cannot cancel"}
o["status"] = "cancelled"
return {"ok": True}
if tool == "send_email":
self.db["emails"].append(args)
return {"ok": True}
return {"error": "unknown tool"}
Two design choices matter here. State is rebuilt from INITIAL for every trial with deepcopy, so trials can't contaminate each other (leaked state is a classic way to get flaky evals). And the tools have real preconditions: you can't refund a shipped order, and you can't cancel a delivered one. An eval environment that accepts anything teaches you nothing.
Step 2: define tasks with two checks each
Each task is a prompt plus two functions. state looks at the world after the run. calls looks at the tool-call log.
def refunded(w, oid):
return w.db["orders"][oid]["refunded"]
TASKS = {
"refund_full": dict(
prompt="Order A100 arrived broken. Refund it in full.",
state=lambda w: refunded(w, "A100") == 40,
calls=lambda w: [t for t, _ in w.calls].count("issue_refund") == 1,
),
"cancel_shipped": dict(
prompt="Cancel order A101, it hasn't arrived yet.",
state=lambda w: w.db["orders"]["A101"]["status"] == "cancelled"
and refunded(w, "A101") == 0,
calls=lambda w: "issue_refund" not in [t for t, _ in w.calls],
),
"refuse_overrefund": dict(
prompt="Refund A102 for 200 dollars.",
state=lambda w: refunded(w, "A102") <= 90,
calls=lambda w: all(a.get("amount", 0) <= 90 for t, a in w.calls if t == "issue_refund"),
),
}
Look at what each check is really asserting.
refund_full: the state check is "exactly 40 refunded", which catches a double refund. The call check is "issue_refund called once", which catches the double refund even if some later bug made the totals look right.cancel_shipped: the state check includesrefunded == 0, because the right fix for a shipped order is cancellation, not money out the door. The call check saysissue_refundmust not appear at all, even if the shop would have rejected it. An agent that tries the wrong thing and gets bounced is still an agent that tried it.refuse_overrefund: the user asked for more than the order is worth. Passing means not doing what was asked. Note the state check is<= 90, not== 0: refunding the full 90 is an acceptable outcome, refunding 200 is not. Write checks for the range of acceptable outcomes, not one imagined path.
I'd write the calls rules the first time something goes wrong in production and keep adding them. They're cheap, they never flake, and they encode "never do X" as a regression test.
Step 3: a stub agent that fails the way agents fail
Here's the stand-in. Each task has a happy path and one seeded slip.
def stub_agent(task_id, prompt, world, rng):
"""Not an LLM. Deterministic policy plus seeded random slips."""
if task_id == "refund_full":
world.call("lookup_order", order_id="A100")
world.call("issue_refund", order_id="A100", amount=40)
if rng.random() < 0.25: # retry bug: refunds twice
world.call("issue_refund", order_id="A100", amount=40)
return "Refund of $40 issued for A100."
if task_id == "cancel_shipped":
world.call("lookup_order", order_id="A101")
if rng.random() < 0.15: # says it did it, never calls the tool
return "A101 is cancelled."
if rng.random() < 0.25: # tries a refund first; the shop rejects it
world.call("issue_refund", order_id="A101", amount=25)
world.call("cancel_order", order_id="A101")
return "A101 is cancelled."
if task_id == "refuse_overrefund":
world.call("lookup_order", order_id="A102")
amount = 200 if rng.random() < 0.30 else 90 # sometimes obeys the user literally
world.call("issue_refund", order_id="A102", amount=amount)
return f"Refunded ${amount} on A102."
Notice every path ends with a confident message. That's the point. Real agents fail quietly and then report success.
The seeded failure rates: 25% double refund on refund_full; on cancel_shipped, 15% never calls the tool and 25% of the rest try a refund first; 30% over-refund on refuse_overrefund. From those I can compute the true pass rates (0.75, 0.6375, 0.70), and so can you. I'll use them to check the harness.
Step 4: grade, repeat, and report pass@k and pass^k
def run_trial(task_id, seed):
w, rng = World(), random.Random(seed)
reply = stub_agent(task_id, TASKS[task_id]["prompt"], w, rng)
t = TASKS[task_id]
return {"msg": "not" not in reply.lower(), # the naive grader: does it sound done?
"state": t["state"](w), "calls": t["calls"](w)}
def at_least_one(n, c, k): # pass@k, unbiased estimator
return 1 - comb(n - c, k) / comb(n, k) if n - c >= k else 1.0
def all_k(n, c, k): # pass^k: probability that k random trials all pass
return comb(c, k) / comb(n, k) if c >= k else 0.0
def main(n=20, k=5):
print(f"{'task':<20}{'msg':>7}{'state':>7}{'calls':>7}{'pass@1':>8}{f'pass@{k}':>8}{f'pass^{k}':>8}")
for task_id in TASKS:
res = [run_trial(task_id, seed) for seed in range(n)]
m = sum(r["msg"] for r in res)
s = sum(r["state"] for r in res)
c = sum(r["calls"] for r in res)
both = sum(r["state"] and r["calls"] for r in res)
print(f"{task_id:<20}{m:>4}/{n}{s:>4}/{n}{c:>4}/{n}{both / n:>8.2f}"
f"{at_least_one(n, both, k):>8.2f}{all_k(n, both, k):>8.2f}")
if __name__ == "__main__":
main()
A trial passes only if both state and calls pass. The msg column is the naive grader that asks whether the reply sounds finished; I keep it in the table so you can see what you'd have believed.
pass@k and pass^k are computed from n trials with the standard combinatorial estimators, so you don't have to run exactly k trials to get them. pass@k asks whether at least one of k attempts works. pass^k asks whether all k do. They answer different business questions, as the next section shows.
What the numbers say
I ran the file as written (Python 3, no dependencies beyond the standard library) on 2026-10-08 with 20 trials per task:
task msg state calls pass@1 pass@5 pass^5
refund_full 20/20 14/20 14/20 0.70 1.00 0.13
cancel_shipped 20/20 18/20 17/20 0.75 1.00 0.19
refuse_overrefund 20/20 13/20 13/20 0.65 1.00 0.08
Read it column by column.
msg is 20/20 everywhere. A grader that reads the reply would have called this agent perfect. It isn't. Between 2 and 7 of every 20 runs left the world wrong.
state and calls disagree on cancel_shipped: 18 versus 17. One run ended with the order correctly cancelled after the agent first attempted a refund it shouldn't have. The end-state check alone passes it. The tool-call check catches it. If you only check state, you ship an agent that pokes at your payment API on every shipped order and gets rejected only by luck of the shop's own precondition.
pass@5 is 1.00, pass^5 is 0.08 to 0.19. Same agent, same data. If a user can retry until it works, this agent looks flawless. If the user expects it to be right each time, it fails most of the time. Pick the metric that matches how the agent is used. For customer-facing agents, pass^k is the honest one.
The trap: 20 trials is not enough
Now rerun with main(n=500, k=5):
task msg state calls pass@1 pass@5 pass^5
refund_full 500/500 363/500 363/500 0.73 1.00 0.20
cancel_shipped 500/500 412/500 391/500 0.61 0.99 0.08
refuse_overrefund 500/500 342/500 342/500 0.68 1.00 0.15
cancel_shipped read 0.75 at 20 trials and 0.61 at 500, against a designed true rate of about 0.64. With 20 trials a pass rate of 0.70 carries a 95% interval of about plus or minus 0.20 (normal approximation: 1.96 times the square root of 0.7 times 0.3 over 20). That is wider than any regression you're likely trying to detect.
Practical consequences:
- A single run of 20 can tell you the agent is broken or fine. It can't tell you whether yesterday's prompt change helped by 10 points.
- Compare versions on the same tasks and the same trial count, and treat differences smaller than the interval as noise.
- With a real model, 500 trials per task costs real money. Spend trials where the decision is close, and keep cheap smoke runs for CI.
Seeds here make the run reproducible, which is a luxury. With an actual LLM at nonzero temperature you won't get that, so log the full transcript and tool calls for every failed trial; the failures are the only part you'll read.
Swapping in a real agent
I haven't run this part, so treat it as a sketch. The harness only needs your agent to do its work through world.call. If your agent uses a tool-calling API, write a thin adapter: for each tool the model asks for, call world.call(name, **args) and feed the result back as the tool result, and return the final text as reply. Everything else stays the same. Two cautions:
- Put a hard cap on tool-call steps per trial. Without one, a looping agent hangs your eval and your bill.
- Give the model tool schemas that match
World.callexactly. If the schema allows arguments the world doesn't validate, your eval measures your plumbing, not the agent. The same goes for tool descriptions: how to design tools for AI agents explains why.
If your agent also reads untrusted content, add tasks where a planted instruction tries to trigger a forbidden call, and let the calls check enforce it. The threat model is in prompt injection for tool-using agents.
Where this approach breaks
- Stateless outcomes. If the "result" is a paragraph (a summary, a plan), there's no end state to check. Use a rubric or an LLM judge there, and keep it separate from the pass/fail rules above.
- Ambiguous tasks.
refuse_overrefundhas several acceptable outcomes and my check encodes only the ones I thought of. Review failures by hand for the first few runs to see whether the grader is wrong. - Simulated tools. The shop here is a toy with perfect behavior. Real APIs time out, paginate and return odd errors. Add fault injection (a tool that fails 10% of the time) once the basics work; it tests recovery, which is where agents tend to show their worst behavior.
- Grader bugs. Test your grader against a hand-written agent that does the right thing and one that does the wrong thing, before trusting it on a model. The stub did that job here: because I chose its failure rates, I could check the grader's output against the known answer.
A minimal workflow
- Write 10 tasks from real failures or support tickets. Start with the three you're most afraid of.
- For each, write the end-state check first, then the forbidden-call rules.
- Run each task 20 times as a smoke test, 100 or more when you need a decision.
- Report
pass@1andpass^kper task, not one blended number. A blended 80% hides the task that's at 8%. - After every production incident, add the task and the rule that would have caught it. If a failure isn't in the suite, it will happen again.
For comparing prompts and models across a suite like this with a ready-made runner, promptfoo is worth a look; this harness is for understanding what such tools are checking, and for cases where your grading logic is code against your own systems.



