What is an AI agent (and why it is not a chatbot) | Rakam AI

AI agent · Updated 2026-08-18

An AI agent is not a chatbot that talks better

The difference is not conversation quality, it is consequence. A conversational assistant produces text. An agent calls tools and changes the state of your system: it creates an entry, sends a message, updates a record. That shift changes everything that matters in production.

What is an AI agent? From text to action

As long as an AI system produces text, an error costs a proofread. As soon as it calls a function that writes to your database, an error costs an accounting correction, an unhappy customer, or an audit finding.

That is why “chatbot or agent?” is not a vocabulary question. It determines:

  • What you measure. For an assistant, the relevance of the answer. For an agent, the rate of correct actions, and the cost of a wrong one.
  • What you authorise. An assistant needs read access. An agent needs write rights, so it needs a perimeter defined action by action.
  • What you keep. An agent must produce a log that holds up in an audit, not just a conversation history.

The four pieces of an agent

An agent usable in production rests on four elements, and missing any one of them makes it undeployable.

The tools. The closed list of functions the agent may call, each with a validated input schema. No arbitrary code execution, no free-form write query. The agent picks from actions you wrote.

The state. Where the record stands, before and after the action. An agent with no notion of state starts over, duplicates, follows up twice. Reading state matters as much as writing it.

The graph. The validated paths the agent is allowed to follow, with the checks required at each step. This is what replaces “the model will figure it out” with “here is what is permitted”.

The threshold. The confidence level below which the agent does nothing alone and hands over. Set per use case, not globally.

Why a graph, rather than freedom

A purely free agent, a model, a list of tools, and “work it out”, works in a demo and fails in production, for a simple reason: on standard cases it finds a correct path, and on rare cases it invents one. Yet rare cases are exactly where a business has rules.

So we describe the activity in a graph: the possible states of a record, the permitted transitions, the mandatory checks before each transition, and the points where a human must validate. The agent does not guess the path, it executes it. That is where the difference lies between a plausible answer and an action you can let through to production.

A concrete example, on bank reconciliation: imported entries pass through counterparty identification, then a confidence score. Above the threshold, posting is automatic. Below it, the record goes to a human validation queue with the proposal and its justification. The full graph is on the system page.

Raising autonomy, rather than granting it

You do not set an agent loose, you raise its autonomy in steps.

ModeWhat the agent doesWhen
On demandThe user asks, the agent executes and shows what it didFrom day one
TriggeredOn event or schedule: the overnight import, the recurring follow-upOnce on-demand mode is stable
DeterministicNo interpretation allowed, fixed path: compliance, invoicing, regulated operationsOnce the workflow is measured
AutonomousThe agent reads the state of the business and acts, under human supervisionOnce confidence is established on a representative volume

A workflow starts out fuzzy. You check it, you score it, and it goes automatic only once it is stable. This progression is the only path we have seen hold at clients who have to answer for their numbers.

What an agent must not do

Three prohibitions we apply systematically, because they are the source of most incidents we have seen elsewhere:

  1. Run a free-form write query. All writes go through functions you wrote, with their checks.
  2. Act without the action being reversible or traced. If an action cannot be undone, it goes through human validation.
  3. Hide its uncertainty. An agent that does not know must say so and hand over, not produce the most likely action.

Agent, RAG, MCP: how they fit together

These three are often conflated although they answer different questions.

  • RAG is how the agent knows: it brings knowledge of your documents and your data to the moment of decision.
  • MCP is how the agent acts: a standard protocol for exposing your business functions as callable tools.
  • The agent is what decides: which tool, in what order, how far without asking.

A serious system has all three. A system with only RAG answers well and does nothing. A system with only tools acts fast and gets it wrong.

Three agents in production, and what they prove

The principles above only count if they hold in operation. Three cases, each on one idea from this page.

Solive, residential property management, on raising autonomy in steps. The day’s bank entries are posted overnight. In the morning the team finds only the cases that need a judgement call in its queue, with the reason for the doubt. The agent does not decide everything: it decides what is clear-cut and documents the rest.

Calixys, financial reconciliation, on the graph rather than freedom. The matching rule is written in plain language, the agent turns it into a deterministic rule and tries it against history before it goes live. What decides in production stays a rule, not a model’s intuition.

VAL Software, business software vendor, on what an agent must not do. The matching engine ranks applications and justifies its ranking criterion by criterion: +17 points of precision, and above all a justification, without which nothing sells in a regulated sector. The data-entry system delivered alongside it carries more than 109 tests, because an agent without tests is not shippable.

Deploying it on your side

Our systems run on your infrastructure, a containerised image deployed on OVH, AWS, GCP or on premises, with local models where sovereignty requires it, including air-gapped networks. Metering is per tenant, which lets software vendors bill usage back to their customers.

Frequently asked

A chatbot produces text, an agent produces effects. A good chatbot explains the procedure for following up on a quote; an agent triggers it: it queries the API to list unapproved quotes, enriches the records, drafts the messages and schedules the sends. The consequence is that the quality criteria change, for a chatbot you judge the relevance of an answer, for an agent you judge the correctness of an irreversible action.

Yes, but never by default and never everywhere. We set a threshold per use case: a workflow starts in proposal mode, every decision is checked and scored, and it only goes automatic once stable across a representative volume. Actions with tax or legal impact stay under double validation permanently. There is no silent execution.

Through your existing APIs and, where they exist, through MCP servers that expose your functions as callable tools. No migration of your software is required and nothing is rewritten: the agent is one more API consumer, with its own rights and its own log.

Every action is logged with the originating question, the tools called, the parameters, the result and the confidence score. That is an AI Act requirement for systems that make decisions, and above all it is what makes the agent usable: without a log, nobody in the company agrees to delegate anything to it.

On a known business activity and with API access, expect six weeks to a production pilot: two weeks of scoping and graph, two weeks of human validation queue and acceptance testing, two weeks of threshold tuning and go live. What we need from you: API access, a set of real documents linked to each other, and one person who knows the business rules.

Newsletter

Every month, what really works in AI for software vendors.

Cases, figures, business models. One email, no more.

Request a demo

Fifteen minutes to see an AI agent working inside software like yours.