What RAG actually solves
A language model does not know your company. It does not know your product references, your internal procedures, or the history of the record your user just opened. Retrieval-augmented generation, or RAG, addresses exactly that gap: at question time, the relevant pieces are retrieved from your data, placed in the model’s context, and the model is asked to answer from those pieces rather than from what it memorised.
The point is not only accuracy. It is traceability: an answer built from identified chunks can cite its sources, so it can be checked. An answer memorised by the model cannot.
The pipeline, step by step
- Ingestion. Your documents come in: PDFs, documentation pages, support tickets, database records. Each one is cleaned and normalised.
- Chunking. Every document is cut into segments. This is the most underestimated step in the pipeline, and we come back to it below.
- Embedding. Each chunk becomes a vector, stored with its metadata: source, date, access rights.
- Retrieval. At question time, the closest chunks are retrieved. Serious retrieval combines vector similarity with keyword search, because the two fail on different cases.
- Reranking. Retrieved chunks are reordered by a finer model that judges actual relevance rather than vector proximity.
- Generation. The model writes the answer from the selected chunks, citing which ones.
The five places it breaks
1. Chunking ignores the document’s structure
This is the number one cause of incomplete answers. Cutting every 500 characters separates a table header from its rows, a condition from its exception, a clause from its sub-clause. The retrieved chunk is then syntactically clean and semantically useless.
What works: chunk along the structure, by section, by clause, by table entry, and keep in each chunk the hierarchical path that situates it. A chunk must be understandable on its own, without the rest of the document.
2. Retrieval is purely vector-based
Vector similarity is excellent on meaning and poor on identifiers. A user looking for reference XRP-4021 or a customer name gets nothing reliable from vector search alone, because those strings have no semantic neighbourhood.
Hybrid retrieval, vector plus lexical, with the two rankings fused, closes that hole for very little engineering. It is almost always the first change that buys the most.
3. There is no evaluation set
Without an evaluation set, every change to the pipeline is a bet. You swap the embedding model, one demo answer improves, and nobody knows what got worse elsewhere.
A useful evaluation set does not need to be large: fifty to two hundred real questions, with the expected answer and the document that contains it. The questions must come from your users, not from an internal writing session, real questions are badly phrased and ambiguous, and that is precisely what needs measuring.
This is why we build evaluation before writing production code. We measure what success means, so we can prove it later.
4. Access rights are applied after retrieval
A frequent and serious design error. If rights filtering happens on the answer rather than on the retrieved chunks, the model has already read documents the user is not allowed to see, and it can restate their content by paraphrase.
Filtering belongs in the retrieval query, against rights metadata attached to each chunk at ingestion.
5. The system always answers
A badly calibrated RAG system has no exit: whatever the question, it produces an answer. On questions your documents cover, that is fine. On the others, it fabricates something plausible, and that answer costs more than no answer, because the user has no way to tell it apart from a good one.
What a usable system does instead: say it does not know, show the closest thing it found, and escalate to a human beyond a threshold set per use case. For actions with tax or legal impact, that threshold drops to zero: human validation every time.
When RAG is no longer enough
RAG answers questions whose answer exists somewhere in a document. It does not answer questions whose answer must be computed, nor those that require following a chain of relations.
“What is the return procedure for a defective item?” is a RAG question. “Why is this invoice blocked, and what would unblock it?” is not: you have to read the record’s state, go back to the order, check the goods receipt, compare against the purchase order. No paragraph contains that answer.
That is where retrieval alone stops and two other pieces take over:
- The business graph, which describes entities, their relations, and the states a record can be in. The question becomes a traversal rather than a text search. This is the subject of our research work on Graph-RAG.
- The agent, which calls tools: query your API, run a request, create an entry. An AI agent does not only search, it acts, a different design problem with different security stakes.
What we do with it for our clients
We deploy these systems onto existing business software, through their APIs, with no migration. RAG is one brick out of four: the business knowledge mined from the software, the customer’s own configuration, the unwritten procedures, and each user’s habits. The detail is on the system page.
Two deployments illustrate the points above better than an explanation.
LinguéO, language learning. The assistant answers the learner citing the original passage, and hands over to a trainer below the confidence threshold. What took it from prototype to production was not a better model: it was its evaluation pipeline, which is exactly the third item on the list above.
École Polytechnique, higher education. Regulations, calendars and administrative procedures, in writing and by voice, in several languages. Academic documents are indexed by passage rather than by document: on a regulation, an approximate answer is worth nothing, and that is the first item on the list.
The SDK we use for ingestion, retrieval and evaluation is published as open source under the name rakam_systems.