What is an AI agent? From text to action
As long as an AI system produces text, an error costs a proofread. As soon as it calls a function that writes to your database, an error costs an accounting correction, an unhappy customer, or an audit finding.
That is why “chatbot or agent?” is not a vocabulary question. It determines:
- What you measure. For an assistant, the relevance of the answer. For an agent, the rate of correct actions, and the cost of a wrong one.
- What you authorise. An assistant needs read access. An agent needs write rights, so it needs a perimeter defined action by action.
- What you keep. An agent must produce a log that holds up in an audit, not just a conversation history.
The four pieces of an agent
An agent usable in production rests on four elements, and missing any one of them makes it undeployable.
The tools. The closed list of functions the agent may call, each with a validated input schema. No arbitrary code execution, no free-form write query. The agent picks from actions you wrote.
The state. Where the record stands, before and after the action. An agent with no notion of state starts over, duplicates, follows up twice. Reading state matters as much as writing it.
The graph. The validated paths the agent is allowed to follow, with the checks required at each step. This is what replaces “the model will figure it out” with “here is what is permitted”.
The threshold. The confidence level below which the agent does nothing alone and hands over. Set per use case, not globally.
Why a graph, rather than freedom
A purely free agent, a model, a list of tools, and “work it out”, works in a demo and fails in production, for a simple reason: on standard cases it finds a correct path, and on rare cases it invents one. Yet rare cases are exactly where a business has rules.
So we describe the activity in a graph: the possible states of a record, the permitted transitions, the mandatory checks before each transition, and the points where a human must validate. The agent does not guess the path, it executes it. That is where the difference lies between a plausible answer and an action you can let through to production.
A concrete example, on bank reconciliation: imported entries pass through counterparty identification, then a confidence score. Above the threshold, posting is automatic. Below it, the record goes to a human validation queue with the proposal and its justification. The full graph is on the system page.
Raising autonomy, rather than granting it
You do not set an agent loose, you raise its autonomy in steps.
| Mode | What the agent does | When |
|---|---|---|
| On demand | The user asks, the agent executes and shows what it did | From day one |
| Triggered | On event or schedule: the overnight import, the recurring follow-up | Once on-demand mode is stable |
| Deterministic | No interpretation allowed, fixed path: compliance, invoicing, regulated operations | Once the workflow is measured |
| Autonomous | The agent reads the state of the business and acts, under human supervision | Once confidence is established on a representative volume |
A workflow starts out fuzzy. You check it, you score it, and it goes automatic only once it is stable. This progression is the only path we have seen hold at clients who have to answer for their numbers.
What an agent must not do
Three prohibitions we apply systematically, because they are the source of most incidents we have seen elsewhere:
- Run a free-form write query. All writes go through functions you wrote, with their checks.
- Act without the action being reversible or traced. If an action cannot be undone, it goes through human validation.
- Hide its uncertainty. An agent that does not know must say so and hand over, not produce the most likely action.
Agent, RAG, MCP: how they fit together
These three are often conflated although they answer different questions.
- RAG is how the agent knows: it brings knowledge of your documents and your data to the moment of decision.
- MCP is how the agent acts: a standard protocol for exposing your business functions as callable tools.
- The agent is what decides: which tool, in what order, how far without asking.
A serious system has all three. A system with only RAG answers well and does nothing. A system with only tools acts fast and gets it wrong.
Three agents in production, and what they prove
The principles above only count if they hold in operation. Three cases, each on one idea from this page.
Solive, residential property management, on raising autonomy in steps. The day’s bank entries are posted overnight. In the morning the team finds only the cases that need a judgement call in its queue, with the reason for the doubt. The agent does not decide everything: it decides what is clear-cut and documents the rest.
Calixys, financial reconciliation, on the graph rather than freedom. The matching rule is written in plain language, the agent turns it into a deterministic rule and tries it against history before it goes live. What decides in production stays a rule, not a model’s intuition.
VAL Software, business software vendor, on what an agent must not do. The matching engine ranks applications and justifies its ranking criterion by criterion: +17 points of precision, and above all a justification, without which nothing sells in a regulated sector. The data-entry system delivered alongside it carries more than 109 tests, because an agent without tests is not shippable.
Deploying it on your side
Our systems run on your infrastructure, a containerised image deployed on OVH, AWS, GCP or on premises, with local models where sovereignty requires it, including air-gapped networks. Metering is per tenant, which lets software vendors bill usage back to their customers.