How to measure an AI agent: 6 indicators | Rakam AI

Measuring an AI agent · Updated 2026-10-06

Measuring an AI agent's performance, before and after go-live

An AI agent is measured on a set of real cases, prepared before it is built, and on six indicators: whether it calls the right tools, faithfulness to sources, consistency, escalation rate, cost per action and the business metric it is meant to move. A successful demo measures none of them.

Why measure an agent differently from a chatbot?

A chatbot is judged on the relevance of an answer. An agent is judged on whether an action was right, because it writes to your software. A wrong action is not fixed by rereading: it is fixed in the database.

So you have to measure what the agent does, not only what it says. And measure it before production, on cases whose right answer is known, then afterwards, on what actually happens.

How do you build an evaluation set?

It is the first step of a project, before the model and before the code. The evaluation set is a list of real cases, each with its expected result, produced with your business experts:

  • common cases, the ones that make up the volume;
  • ambiguous cases, the ones that decide whether the agent is usable;
  • out-of-scope cases, where the right answer is to refuse or hand over.

The test set is kept separate from whatever is used to configure the agent. Otherwise you measure its memory, not its accuracy. The full method is in how to build an AI agent.

Which indicators should you track?

Six indicators cover the essentials. The first four are measured on the evaluation set, the last two in production.

IndicatorWhat it measures
Tool-call accuracyDid the agent call the right tool, with the right parameters?
FaithfulnessIs the answer grounded in the sources, with nothing invented?
ConsistencyDoes the same question get the same answer?
Escalation rateWhat share of requests does the agent pass to a human, and rightly so?
Cost per actionWhat does a successful action cost, model calls included?
Business metricWhat the agent is meant to move: tickets, lead times, hours, revenue

The escalation rate reads both ways. Too low, and the agent takes decisions it should not. Too high, and it does not deliver the expected service. The target is set per use, on the cost of a mistake.

What do these measurements look like on a delivered project?

At OOTI, which builds an ERP for architecture firms, the agent turns natural-language questions into queries on the ERP data. The published measurements:

  • 92% tool-call accuracy;
  • 97% faithfulness score;
  • 100% response consistency;
  • under $0.10 per question.

Each figure answers a different question. Tool accuracy says whether the agent queries the right data. Faithfulness says whether it reports what it found without embellishing it. Consistency says that a partner and a project manager asking the same question get the same figure. Cost says whether usage holds up at scale. The details are in the case study.

How do you compare two configurations?

On the same test set, with the same protocol, publishing the gap. At VAL Software, CV-job matching gained 17 points of accuracy by comparing 28 configurations on the same test set. Without a common protocol, you would be comparing impressions.

The system ships with more than 109 tests, which makes it possible to check that an improvement in one place has not broken anything elsewhere.

What should you measure once in production?

What the evaluation set does not see: the questions nobody anticipated, drift when your software changes, the real cost at scale. The log of every action is the raw material here. A mistake found in production becomes one more case in the evaluation set, and the next measurement covers it.

And the business metric, the one that justifies the price. It is what lets you bill on outcome rather than on licence, as the credit model explains. If you want to know where to start on your own software, the scan below gives a first diagnosis.

Frequently asked

Enough to cover the real families of questions, including the ambiguous cases and those where the agent should refuse or hand over. The number depends on how varied the activity is. What matters more than volume: cases produced by your business experts, with the expected result written down before the test.

Yes, for faithfulness and relevance, provided you check the grading against a hand-graded sample. For tool-call accuracy, a direct comparison with the expected call is more reliable: the right tool, with the right parameters, or not.

At every change: of model, prompt, graph, or version of your software. An automated test set reruns in minutes. That is what lets you switch models without fear, and prove a change has not degraded anything elsewhere.

The business metric, with how it was measured. A customer does not buy tool accuracy, they buy tickets avoided, records processed, hours given back. The technical indicators are there to guarantee that figure holds over time.

Newsletter

Every month, what really works in AI for software vendors.

Cases, figures, business models. One email, no more.

Request a demo

Fifteen minutes to see an AI agent working inside software like yours.