What is RAG? An AI assistant that works with company documents

A wooden library card catalogue with one drawer open, a card on the desk and a brass magnifying glass

RAG (retrieval-augmented generation) is a method in which an AI model searches your documents before answering a question, and bases its answer only on the passages it finds. The model does not memorise your company’s knowledge; for each question it finds the relevant documents, reads them, and answers with its sources shown. Most in-house AI assistants are built this way.

Below we explain how RAG works, why it is effective, the steps we follow when building a company assistant, and the pitfalls teams fall into most often.

What is RAG? An analogy

A large language model is like someone who has read widely but has never met your company. They answer general questions well. Ask them, “How many days do business customers have under our returns policy?” and they will either admit they do not know or — worse — invent an answer that sounds plausible.

RAG is asking that person to go to the company archive and put the relevant files on the desk before answering. The answer now comes not from memory but from the documents in front of them. And they can show you which document they looked at.

How does RAG work? Four steps

A RAG system works in two stages: preparation first, then retrieval and answering for each question.

1. Splitting documents into chunks

Contracts, procedures, product manuals and support tickets are split into meaningful chunks — usually a few paragraphs each. Alongside every chunk goes a note of which document, which section and which date it came from.

2. Making the chunks searchable

Each chunk is converted into a numerical vector that represents its meaning, and stored in a vector database. This means search works on meaning, not just matching words. Someone asking about a “refund” can find a document that only says “reimbursement”.

3. Finding the best chunks for the question

When a user asks something, the system converts the question into a vector in the same way and retrieves the closest chunks. Good systems add classic keyword search and a re-ranking step at this point.

4. Grounding the answer in the chunks

The retrieved chunks go to the model together with the question, along with a clear instruction: “Answer only from these texts. If the answer is not in them, say you don’t know. State which document you relied on.”

Why RAG rather than fine-tuning?

Another way to teach a model your company’s knowledge is fine-tuning — retraining the model on your own data. For most company assistants, RAG is the better choice:

RAGFine-tuning
When a document changesAdd the new document; it is used at onceThe model must be retrained
Citing sourcesComes naturallyDifficult
Access controlCan be applied per documentImpractical
What it is good atKnowledge-based answersTone, format, a narrow task

Fine-tuning is good at changing how a model speaks. RAG decides what it speaks from. The two can be combined, but a knowledge problem is usually solved with RAG.

An in-house AI assistant: a concrete scenario

Picture a fifty-person company. HR policies, purchasing procedures, product technical documents and years of support tickets sit in scattered folders. A new joiner asks three people to answer one simple question.

With a RAG assistant in place:

  • Someone in sales asks, “Which parts are not covered by the warranty on product X?” The assistant finds the relevant section of the manual, writes the answer, and shows the document name and section.
  • An employee asks, “How does carrying over annual leave work?” The assistant answers from the current HR policy; the old version stays in the archive and is excluded from search.
  • Someone asks about salary tables. The assistant has no permission for those documents, so it does not answer — it points them to someone who does.

That last point matters. A good company assistant never goes beyond the documents the user can already access.

The steps we follow when building one

Choosing the sources

Not every document belongs in the assistant. Outdated, contradictory or draft documents lower the quality of answers. The first job is deciding which sources count as “correct”. That is usually an organisational decision, not a technical one.

Modelling permissions

Which user may see which document? That information is carried into the search stage. The model never sees a chunk the user is not allowed to see.

Deciding where data is processed

Which model the documents go to, where they are stored and how long logs are kept are set out in writing at the start. For companies handling sensitive data — and with obligations under GDPR or, in Türkiye, the similar KVKK — this is the first item in the architecture.

Measuring

Before going live, we prepare a test set from real questions: the questions, the expected answers and the documents each answer should rest on. The set is run again after every change. Without measurement, you cannot honestly say “it works well”.

Starting small

We open the first version to one team and one group of documents — for example, only the support team and the product manuals. The team uses the assistant in its real work, and we see the questions it cannot answer, the documents people find wrong and the topics asked about most. That feedback is the most valuable input for improving chunking and search before the second document group is added. An assistant opened to everyone at once is far harder to debug.

The most common pitfalls

  • Poor chunking. Splitting a table down the middle, or separating a heading from its content, produces wrong answers even when the right document is found.
  • Old documents. If three versions of the same policy are searchable, the model cannot tell which one applies.
  • No sources shown. Users do not trust an answer whose source they cannot see — nor should they.
  • Unable to say “I don’t know”. Unless the model is told clearly to say so, it fills the gap by making something up. This is the problem known as hallucination, which we cover separately in AI hallucination, privacy and human approval.
  • No document owners. If no one is responsible for keeping each document group current, the assistant gradually starts answering from stale information.
  • Not measuring search quality. Most errors come from search, not from the model. If the right chunk is not found, even the best model cannot give the right answer.

Beyond RAG: when the assistant starts doing things

A RAG assistant reads and answers. The next step is for the assistant to act inside company systems — opening a ticket, updating a record, producing a report. At that point it becomes an AI agent, and questions of permission, approval and logging matter even more.

Frequently asked questions

Do we need to train a model on our documents for RAG?

No — that is RAG’s main advantage. The model is not retrained; the documents are made searchable, and the relevant chunks are given to the model with each question. When a new document is added, the assistant can use it shortly afterwards.

Will the assistant never give a wrong answer?

It can. RAG lowers the rate of wrong answers but does not eliminate it. That is why we insist that answers show their sources, that the assistant says when it does not know, and that it is measured against a regular test set.

What kinds of documents does it work with?

Almost any document that contains text: contracts, procedures, manuals, emails, support tickets. Scanned documents must first be converted to text. Tables and images need special preparation.

What determines the cost of a RAG assistant?

Document volume, the number of questions, the length of answers and the model used are the main factors. The variety of documents, the permission structure and how often things are updated shape the scope of the set-up, and scanned documents or complex tables need extra preparation. We settle the scope together in a first conversation and follow it with a written proposal.

Will confidential documents be visible to others?

Not in a properly built system. Access control is applied at the search stage, so a document the user cannot access never reaches the model.

If you would like to talk about an assistant that works with your own company documents, see our AI systems architecture page or write to us.

Open a conversation