Your AI Agent Passed Evals and Still Failed in Production: Building an Eval Harness That Matches Reality
Why AI agents pass offline evals and still fail customers, and how to build an AI agent evaluation harness from real traces, calibrated judges, and CI gates.
If your agent passed evals and still failed a customer, your eval set was measuring the wrong world. Offline test cases drift away from production because of distribution shift, tool side effects that mocks hide, and multi-turn conversations that single-turn tests never touch. The fix is an AI agent evaluation harness built from real production traces, scored by calibrated judges, and enforced in CI.
You are not alone in this. VentureBeat’s VB Pulse survey, fielded in June 2026 across 157 organizations with 100 or more employees, found that 50% had shipped an AI feature that cleared internal evaluations and then caused a customer-facing failure. Only 5% said they fully trust automated evaluation as it stands, and the most-cited limitation (29%) was that evaluations align poorly with real-world outcomes. VentureBeat itself calls the sample self-selected and directional, so treat the numbers as a signal, not a census. The signal is loud enough.
Why do offline evals diverge from production?
An eval suite is a model of your users. Like any model, it is wrong in specific, predictable ways. Here are the three that bite agents hardest.
Distribution shift
Your golden set was written by engineers during development. Customers write differently: typos, half-sentences, two requests in one message, Arabic and English mixed in the same turn (a very normal pattern for GCC users). They also ask about things that did not exist when the set was written, like a new product, a new policy, or a new error message. Your pass rate stays green because nobody is testing the inputs that actually arrive.
Tool-call side effects
Most agent evals mock the tools. The mock returns a clean JSON payload in 50ms, every time. In production the CRM API times out, returns a paginated list, or succeeds on the second retry after creating a duplicate record. A text-only eval checks whether the final answer sounds right. It cannot see that the agent issued two refunds to get there. This is why we push teams toward trajectory evaluation: score the sequence of tool calls and their arguments, not just the last message.
Multi-turn drift
Single-turn test cases ask one question and grade one answer. Real sessions run 8 to 20 turns, with the user correcting the agent, changing their mind, or pasting in a document halfway through. Context windows fill up, earlier instructions get crowded out, and the agent forgets a constraint it respected in turn two. If your harness never runs a long conversation, it will never catch this class of failure.
What does an eval harness that matches reality look like?
An AI agent evaluation harness has five parts. Skip one and the gap reopens.
| Layer | What it does | Typical tooling |
|---|---|---|
| Trace capture | Records every prompt, tool call, argument, result, latency, and cost in production | Langfuse, Arize Phoenix, Braintrust |
| Golden set | Versioned cases built from real traces plus known incidents | Datasets in your tracing tool, or plain JSONL in git |
| Scorers | Deterministic checks plus calibrated LLM judges | DeepEval, Ragas, custom pytest assertions |
| CI gates | Blocks a prompt, model, or tool change that regresses the golden set | GitHub Actions or GitLab CI running the scorers |
| Online monitoring | Samples live traffic and scores it continuously | Same tracing platform, with sampled judge runs |
The arrows matter more than the boxes. Traces feed the golden set, the golden set feeds CI, and online monitoring feeds new failures back into traces. That loop is the whole point.
How do you turn production traces into a golden set?
Start with what already went wrong. Here is the process we use:
- Pull 2 to 4 weeks of traces and cluster them by intent (a quick embedding plus clustering pass works, or just tag by the first tool called).
- Sample from every cluster, not just the big ones. Rare intents are where agents improvise.
- Add every known incident as a permanent case: the support escalation, the wrong refund, the hallucinated policy.
- Write the expected outcome, not the expected wording. For agents that means the expected end state (ticket created, refund not issued) and the acceptable tool sequences.
- Keep full multi-turn sessions as test cases, and replay them turn by turn.
- Version the set in git and record which version each release was tested against.
Scrub personal data before traces leave production. For UAE and KSA teams, check the set against your PDPL obligations before you hand it to any SaaS eval tool.
How do you calibrate an LLM-as-judge?
An LLM judge is a measuring instrument. You would not trust a thermometer you never checked against a known temperature.
- Label a sample by hand. Have two people mark a few hundred real outputs pass or fail against the rubric.
- Run the judge on the same sample and measure agreement with the humans. Look at the disagreements, not just the score.
- Split vague rubrics into narrow yes/no questions. “Was the answer helpful?” drifts. “Did the agent confirm before cancelling the order?” does not.
- Pin the judge model and prompt version. Change either one and you have a new instrument with a new baseline.
- Re-check on a schedule, because your traffic changes even if the judge does not.
Use deterministic checks wherever you can: tool name matches, arguments validate against the schema, no write call without a preceding confirmation. Save the judge for what code cannot check.
Which tools fit which part of the harness?
| Tool | Best at | License / status (Oct 2026) | Watch out for |
|---|---|---|---|
| Braintrust | Experiments, dataset management, LLM-judge scorers, tracing | Proprietary SaaS; raised an $80M Series B at an $800M valuation in Feb 2026 | Vendor lock-in on datasets and scorers |
| Langfuse | Open tracing, datasets, prompt management, evals | MIT-licensed core; acquired by ClickHouse in Jan 2026 | Self-hosting means you run ClickHouse too |
| DeepEval | pytest-style CI gates; agent metrics like task completion and tool correctness | Apache 2.0 | Judge-based metrics still need calibration |
| Ragas | Retrieval metrics plus agent metrics such as tool call accuracy and goal accuracy | Apache 2.0 | Reference-based tool metrics are strict about call order by default |
| Promptfoo | YAML test suites, red-teaming, multi-model comparison | MIT; OpenAI announced its acquisition on 9 March 2026 and is folding it into OpenAI Frontier | Neutrality questions for teams comparing model vendors |
| Arize Phoenix | OpenTelemetry-native tracing with built-in eval primitives | Source-available, free to self-host | Fewer turnkey CI patterns than DeepEval |
For deeper head-to-heads, see our DeepEval vs RAGAS comparison and the Promptfoo alternatives guide if the OpenAI ownership change is a concern for your procurement team.
The practical default for most teams: one tracing platform (Langfuse or Phoenix if you want open source, Braintrust if you want managed) plus DeepEval for CI gates and Ragas where retrieval quality matters.
We build an AI agent evaluation harness from your production traces: golden set, calibrated judges, CI gates, and online monitoring, on the tools you already use.
Set up my eval harnessHow should CI gates work for agents?
Treat the golden set like a unit test suite, with two differences: results are probabilistic, and some failures matter far more than others.
- Run each case several times and gate on pass rate, not a single run. One successful run proves the agent can do the task, not that it will.
- Tier your cases. A safety case (never refund without confirmation) is a hard block at any failure. A tone case can tolerate small dips.
- Gate on deltas, not absolutes. “No more than 2 points below the last release” is easier to live with than a fixed threshold that was set before your traffic changed.
- Trigger on every change that can alter behavior: system prompt, model version, tool schema, retrieval index. Model upgrades are the most common silent regression.
- Track cost and latency in the same run. An agent that passes by making 14 tool calls instead of 3 is a regression too.
What should online monitoring catch?
The VentureBeat survey found that only 23% of organizations monitor whether their agent’s answers are right, while most monitor only whether it is running. Uptime dashboards will not tell you the agent started inventing return policies on Tuesday.
At minimum, sample a slice of live sessions and run the same calibrated judges you use in CI. Alert on drops in task success, spikes in tool errors or retries, and new intent clusters you have no golden cases for. Then close the loop: every flagged session goes to a human for triage, and confirmed failures go into the golden set. Within a few weeks, your eval set starts to look like your production traffic, which was the goal all along.
For agent-specific scoring beyond final answers, our agent trajectory testing guide covers how to score tool sequences step by step.
Where to start this week
Pick your ten worst production conversations from last month. Turn them into test cases with expected end states. Run your current agent against them. If it fails more than one or two, you have just measured your evaluation gap, and you know exactly which cases the next release must pass.
When you want help building the full loop, our Agent Trajectory Testing Sprint is a fixed-scope, 5-day engagement: we curate golden trajectories from your traces, deploy the harness on your stack, calibrate the judges, and wire the gates into CI. Talk to us about your agent.
Frequently Asked Questions
Why do AI agents pass evals but fail in production?
Because the eval set only covers the cases someone thought to write. Real users bring different phrasing, longer conversations, and messier data, and real tools have side effects that mocked tools hide. The fix is an AI agent evaluation harness fed by production traces, so the test set keeps moving toward what customers actually do instead of what the team imagined.
What is an AI agent evaluation harness?
It is the plumbing that turns agent quality into a repeatable number: a versioned golden dataset, scorers (deterministic checks plus calibrated LLM judges), a runner that replays agent trajectories including tool calls, CI gates that block regressions, and online monitoring that samples live traffic. Tools like DeepEval, Ragas, Langfuse, Braintrust, and Arize Phoenix supply parts of it; the eval harness is how you wire them together.
How do you calibrate an LLM-as-judge for agent evals?
Have humans label a sample of real outputs pass or fail, run the judge on the same sample, and measure agreement. Where they disagree, tighten the rubric or split it into narrower yes/no questions, then re-measure. Pin the judge model version, because swapping it silently changes your baseline. An uncalibrated LLM judge is just a second opinion you have not checked.
Which tools should I use to build an agent eval harness in 2026?
Most teams combine one tracing or observability platform (Langfuse, Arize Phoenix, or Braintrust) with one metrics library (DeepEval for pytest-style CI gates, Ragas for retrieval and agent goal metrics). Promptfoo is still MIT-licensed after OpenAI agreed to acquire it, but vendor-neutral teams are weighing that. The tool matters less than having production traces flowing into the golden set.
Should automated evals be allowed to ship agent changes on their own?
Only for low-risk changes, and only once your judges are calibrated and your golden set includes real incidents. Until then, use CI eval gates to block obvious regressions and keep a human sign-off for anything touching money, data deletion, or customer communication. Autonomy should follow measured trust in the harness, not precede it.
Complementary NomadX Services
Related Articles
Break It Before They Do.
Book a free 30-minute GenAI QA scope call. We review your AI application, identify the top risks, and show you exactly what to test before you ship.
Every engagement is scoped by our principal architect, Adrian Vale: 20+ years in production engineering, 40+ professional certifications. Meet Adrian
Talk to an Expert