RAG: how it works, and where it breaks | Rakam AI

RAG · Updated 2026-08-18

RAG, and the five places it breaks in production

Retrieval-augmented generation has become the default building block of every enterprise AI project. It is also the most badly calibrated: a prototype takes three days, a version that holds in production means solving five problems the prototype hides.

What RAG actually solves

A language model does not know your company. It does not know your product references, your internal procedures, or the history of the record your user just opened. Retrieval-augmented generation, or RAG, addresses exactly that gap: at question time, the relevant pieces are retrieved from your data, placed in the model’s context, and the model is asked to answer from those pieces rather than from what it memorised.

The point is not only accuracy. It is traceability: an answer built from identified chunks can cite its sources, so it can be checked. An answer memorised by the model cannot.

The pipeline, step by step

  1. Ingestion. Your documents come in: PDFs, documentation pages, support tickets, database records. Each one is cleaned and normalised.
  2. Chunking. Every document is cut into segments. This is the most underestimated step in the pipeline, and we come back to it below.
  3. Embedding. Each chunk becomes a vector, stored with its metadata: source, date, access rights.
  4. Retrieval. At question time, the closest chunks are retrieved. Serious retrieval combines vector similarity with keyword search, because the two fail on different cases.
  5. Reranking. Retrieved chunks are reordered by a finer model that judges actual relevance rather than vector proximity.
  6. Generation. The model writes the answer from the selected chunks, citing which ones.

The five places it breaks

1. Chunking ignores the document’s structure

This is the number one cause of incomplete answers. Cutting every 500 characters separates a table header from its rows, a condition from its exception, a clause from its sub-clause. The retrieved chunk is then syntactically clean and semantically useless.

What works: chunk along the structure, by section, by clause, by table entry, and keep in each chunk the hierarchical path that situates it. A chunk must be understandable on its own, without the rest of the document.

2. Retrieval is purely vector-based

Vector similarity is excellent on meaning and poor on identifiers. A user looking for reference XRP-4021 or a customer name gets nothing reliable from vector search alone, because those strings have no semantic neighbourhood.

Hybrid retrieval, vector plus lexical, with the two rankings fused, closes that hole for very little engineering. It is almost always the first change that buys the most.

3. There is no evaluation set

Without an evaluation set, every change to the pipeline is a bet. You swap the embedding model, one demo answer improves, and nobody knows what got worse elsewhere.

A useful evaluation set does not need to be large: fifty to two hundred real questions, with the expected answer and the document that contains it. The questions must come from your users, not from an internal writing session, real questions are badly phrased and ambiguous, and that is precisely what needs measuring.

This is why we build evaluation before writing production code. We measure what success means, so we can prove it later.

4. Access rights are applied after retrieval

A frequent and serious design error. If rights filtering happens on the answer rather than on the retrieved chunks, the model has already read documents the user is not allowed to see, and it can restate their content by paraphrase.

Filtering belongs in the retrieval query, against rights metadata attached to each chunk at ingestion.

5. The system always answers

A badly calibrated RAG system has no exit: whatever the question, it produces an answer. On questions your documents cover, that is fine. On the others, it fabricates something plausible, and that answer costs more than no answer, because the user has no way to tell it apart from a good one.

What a usable system does instead: say it does not know, show the closest thing it found, and escalate to a human beyond a threshold set per use case. For actions with tax or legal impact, that threshold drops to zero: human validation every time.

When RAG is no longer enough

RAG answers questions whose answer exists somewhere in a document. It does not answer questions whose answer must be computed, nor those that require following a chain of relations.

“What is the return procedure for a defective item?” is a RAG question. “Why is this invoice blocked, and what would unblock it?” is not: you have to read the record’s state, go back to the order, check the goods receipt, compare against the purchase order. No paragraph contains that answer.

That is where retrieval alone stops and two other pieces take over:

  • The business graph, which describes entities, their relations, and the states a record can be in. The question becomes a traversal rather than a text search. This is the subject of our research work on Graph-RAG.
  • The agent, which calls tools: query your API, run a request, create an entry. An AI agent does not only search, it acts, a different design problem with different security stakes.

What we do with it for our clients

We deploy these systems onto existing business software, through their APIs, with no migration. RAG is one brick out of four: the business knowledge mined from the software, the customer’s own configuration, the unwritten procedures, and each user’s habits. The detail is on the system page.

Two deployments illustrate the points above better than an explanation.

LinguéO, language learning. The assistant answers the learner citing the original passage, and hands over to a trainer below the confidence threshold. What took it from prototype to production was not a better model: it was its evaluation pipeline, which is exactly the third item on the list above.

École Polytechnique, higher education. Regulations, calendars and administrative procedures, in writing and by voice, in several languages. Academic documents are indexed by passage rather than by document: on a regulation, an approximate answer is worth nothing, and that is the first item on the list.

The SDK we use for ingestion, retrieval and evaluation is published as open source under the name rakam_systems.

Frequently asked

They are not alternatives, they answer different problems. RAG solves a knowledge problem: the model does not know your data, so you bring it in at question time. Fine-tuning solves a behaviour problem: the model should answer in a format, tone or structure you want to impose. A need for up-to-date knowledge is handled with RAG, because retraining on every documentation change is untenable. A need for strict formatting is handled with fine-tuning or structured output. Many serious systems do both.

A demonstrable prototype takes days. A production version, with evaluation, source traceability, update handling and quality monitoring, takes weeks. The gap between the two is not more code: it is the evaluation set, a chunking strategy matched to your documents, and the decision about what the system does when it does not know.

That question decides whether a RAG system is usable at all. A system that always answers produces plausible, wrong answers, which cost more than no answer. Our systems are configured to say they do not know, cite the closest thing they found, and escalate to a human beyond a confidence threshold set per use case.

Not always. Below a few hundred thousand chunks, a vector extension on the database you already run, PostgreSQL with pgvector, for instance, is plenty and saves you a component to operate. A dedicated store is justified by large volumes, complex metadata filtering, or tight latency constraints.

That depends entirely on the architecture, and it is a choice rather than a given. We deploy systems on your infrastructure, with local models where sovereignty requires it, including air-gapped networks. When a remote model is used, personal data obfuscation happens before the call.

Newsletter

Every month, what really works in AI for software vendors.

Cases, figures, business models. One email, no more.

Request a demo

Fifteen minutes to see an AI agent working inside software like yours.