Why measure an agent differently from a chatbot?
A chatbot is judged on the relevance of an answer. An agent is judged on whether an action was right, because it writes to your software. A wrong action is not fixed by rereading: it is fixed in the database.
So you have to measure what the agent does, not only what it says. And measure it before production, on cases whose right answer is known, then afterwards, on what actually happens.
How do you build an evaluation set?
It is the first step of a project, before the model and before the code. The evaluation set is a list of real cases, each with its expected result, produced with your business experts:
- common cases, the ones that make up the volume;
- ambiguous cases, the ones that decide whether the agent is usable;
- out-of-scope cases, where the right answer is to refuse or hand over.
The test set is kept separate from whatever is used to configure the agent. Otherwise you measure its memory, not its accuracy. The full method is in how to build an AI agent.
Which indicators should you track?
Six indicators cover the essentials. The first four are measured on the evaluation set, the last two in production.
| Indicator | What it measures |
|---|---|
| Tool-call accuracy | Did the agent call the right tool, with the right parameters? |
| Faithfulness | Is the answer grounded in the sources, with nothing invented? |
| Consistency | Does the same question get the same answer? |
| Escalation rate | What share of requests does the agent pass to a human, and rightly so? |
| Cost per action | What does a successful action cost, model calls included? |
| Business metric | What the agent is meant to move: tickets, lead times, hours, revenue |
The escalation rate reads both ways. Too low, and the agent takes decisions it should not. Too high, and it does not deliver the expected service. The target is set per use, on the cost of a mistake.
What do these measurements look like on a delivered project?
At OOTI, which builds an ERP for architecture firms, the agent turns natural-language questions into queries on the ERP data. The published measurements:
- 92% tool-call accuracy;
- 97% faithfulness score;
- 100% response consistency;
- under $0.10 per question.
Each figure answers a different question. Tool accuracy says whether the agent queries the right data. Faithfulness says whether it reports what it found without embellishing it. Consistency says that a partner and a project manager asking the same question get the same figure. Cost says whether usage holds up at scale. The details are in the case study.
How do you compare two configurations?
On the same test set, with the same protocol, publishing the gap. At VAL Software, CV-job matching gained 17 points of accuracy by comparing 28 configurations on the same test set. Without a common protocol, you would be comparing impressions.
The system ships with more than 109 tests, which makes it possible to check that an improvement in one place has not broken anything elsewhere.
What should you measure once in production?
What the evaluation set does not see: the questions nobody anticipated, drift when your software changes, the real cost at scale. The log of every action is the raw material here. A mistake found in production becomes one more case in the evaluation set, and the next measurement covers it.
And the business metric, the one that justifies the price. It is what lets you bill on outcome rather than on licence, as the credit model explains. If you want to know where to start on your own software, the scan below gives a first diagnosis.