All articles
Agentic AI 8 min read ·

How to Actually Evaluate an Agent

'It worked when I tried it' is not evaluation. Here is how to measure an agent that takes a different path every time.

By NeuralNetworki.ng Team · AI Engineers

Why agents are hard to evaluate

Evaluating ordinary software is straightforward because it is deterministic: the same input yields the same output, so a fixed set of test cases pins down its behaviour. Agents break that assumption in both directions. The same input can produce different paths on different runs, several of which are perfectly correct. And the same path can be correct one day and subtly wrong the next, because a tool returned something slightly different or the model's reasoning drifted. You cannot capture that with a handful of unit tests, and teams that try are repeatedly blindsided by behaviour their tests never covered.

So evaluating an agent is less like testing a function and more like assessing a new employee: you care about results, but you also care about how they got there, and you judge them over many tasks, not one.

Evaluate outcomes, then trajectories

Start where users start: the outcome. Did the agent achieve the goal? For tasks with a checkable result, did it book the right slot, return the correct figure, file the complaint to the right department, this is the score that matters most, and it should anchor your evaluation.

But outcome alone is not enough, because two agents can reach the same correct answer in wildly different ways. One might do it in three clean steps; the other might flail through fifteen, call an expensive tool repeatedly, and stumble into the answer by luck. So also evaluate the trajectory: how many steps it took, which tools it used, how much it cost, how long it ran, and crucially whether any step along the way was unsafe or wasteful. An agent that gets the right answer through a reckless path is a production incident waiting to happen, even if the final output looks fine.

Build a test set of real tasks

You cannot evaluate against your imagination. Collect a set of real goals the agent will actually face, drawn from real users where possible, and where you can, record a known-good outcome for each. This test set is the foundation of everything; invest in it.

Then, and this is the part people skip, run the agent across the whole set repeatedly, not once. Variance is not noise to be ignored; it is a property of the system you are measuring. A single pass that happens to succeed tells you almost nothing about reliability. Running each task many times gives you a distribution: a success rate, a spread of costs, a sense of the worst case. "It passes 96% of the time, and when it fails it fails safely" is a statement you can ship on. "It worked when I tried it" is not.

Use judges, but verify the judge

Many agent outputs are open-ended, a written summary, a plan, a response, and have no single correct string to compare against. Here a second model acting as a judge can score quality at scale, which is the only practical way to evaluate thousands of runs.

But treat the judge as a noisy instrument, not an oracle. Calibrate it: have humans rate a sample, and check that the judge's scores correlate with theirs before trusting it at scale. And frame its task carefully. A judge asked "is this good?" tends toward generous, vague approval. A judge asked to actively find faults, "list every way this output is wrong, incomplete, or unsafe", is far more useful, because criticism surfaces problems that a thumbs-up hides. For high-stakes judgements, use several judges with different framings and look for agreement.

Track regressions over time

An agent is never finished; you will change its prompt, swap its model, add or modify tools. Every one of those changes can silently make it worse on cases it used to handle. The model that is better on average might be worse on your specific task; the prompt tweak that fixed one failure might have broken three others.

The defence is to run your evaluation suite on every meaningful change and compare against the previous baseline, exactly as you would run a regression test suite for ordinary code. An agent that quietly degraded after a "small" prompt edit is one of the most common, and most avoidable, production surprises. Make the eval part of the change process, not something you run once before launch and forget.

Measure cost and latency as first-class metrics

Finally, remember that quality is not the only axis users care about. An agent that produces excellent results but takes thirty seconds and costs a dollar per request may be unusable for your product, while a slightly less capable one that responds in two seconds for a cent is a hit. Cost-per-task and time-to-completion are not afterthoughts; they belong right next to accuracy on your evaluation dashboard.

Putting them side by side forces the trade-offs into the open. Maybe a cheaper model handles 90% of cases and you escalate the rest. Maybe you cap the loop tighter and accept a slightly lower success rate for half the cost. These are product decisions, and you can only make them deliberately if you are measuring all three, quality, cost, and speed, together.

#Agentic AI#Evaluation#Testing

Related work

This is the kind of problem we solve in Agentic AI Systems. See it in practice in our Agentic Honeypot, ARGUS case study.

Talk to us about your project