What Is Retrieval-Augmented Generation (RAG)? A Plain-English Guide

In short

Retrieval-augmented generation (RAG) is a technique where an AI system first searches a trusted collection of documents for passages relevant to your question, then hands those passages to a language model along with the question — so the answer is grounded in real sources instead of the model's memory alone.

Ask a chatbot about your company's refund policy, last week's news, or a document it has never seen, and one of two things happens: it admits it doesn't know, or it confidently makes something up. Retrieval-augmented generation — RAG — is the most common fix, and it's behind a large share of the "AI assistant" features you'll meet at work and online.

The idea is simpler than the name. Here it is without the jargon.

The problem RAG solves

A large language model learns from a huge snapshot of text, and then training stops. After that point it knows nothing new, and it never knew anything that wasn't in the snapshot — your internal wiki, your product manual, yesterday's pricing change.

When a model is asked about something outside that snapshot, it doesn't have a reliable way to say "that's not in my training data". It predicts plausible-sounding text, and plausible is not the same as true. These confident wrong answers are usually called hallucinations. Retraining a model every time your documents change would be slow and expensive, so RAG takes a different route: leave the model alone and feed it the facts at the moment it needs them.

How it works, in three steps

Every RAG system, from a weekend project to an enterprise assistant, follows the same basic shape.

  • Index: split your documents into short chunks and store them in a search index. Most systems also convert each chunk into an embedding — a list of numbers that captures its meaning — so passages can be found by meaning, not just matching words.
  • Retrieve: when a question arrives, search the index for the chunks most relevant to it. Many systems combine meaning-based search with ordinary keyword search, because each catches things the other misses.
  • Generate: place the best few chunks into the prompt alongside the question, with an instruction to answer using only those sources — and ideally to cite them, and to say so when they don't cover the question.

An everyday analogy

Think of the difference between a closed-book and an open-book exam. A plain language model is sitting the closed-book version: it can only write what it remembers. RAG turns it into an open-book exam, with a helpful assistant who finds the right pages and puts them on the desk before each question.

The student hasn't got any smarter. They've just stopped relying on memory for facts they can look up — which is exactly how you'd want a careful person to work, too.

Why it reduces made-up answers — but doesn't eliminate them

Grounding answers in retrieved text makes them more accurate and, just as importantly, checkable: a good RAG system shows you the passage it relied on, so you can verify the claim in seconds.

But RAG moves the weak point rather than removing it. If the search step fetches the wrong passage, the model will answer confidently from the wrong passage. If the documents are outdated, so is the answer. And models can still misread a source or blend it with what they already "know". The quality of a RAG system is mostly the quality of its retrieval and its documents.

RAG vs fine-tuning

Fine-tuning means continuing to train a model on your own examples, which changes the model itself. It's good at teaching consistent behaviour — a tone of voice, an output format, a narrow task done the same way every time. It's a poor way to teach facts that change, because every change means training again.

RAG does the opposite: it leaves the model unchanged and supplies facts at question time, so updating knowledge is as simple as updating a document. A useful rule of thumb is RAG for knowledge, fine-tuning for behaviour. Plenty of real systems use both.

What separates a good RAG system from a bad one

If you're evaluating an AI assistant that claims to answer from your documents, these are the things worth asking about.

  • Clean, current sources — duplicated or out-of-date documents produce contradictory answers
  • Sensible chunking — passages long enough to keep their context, short enough to stay on topic
  • Good retrieval — usually a mix of keyword and meaning-based search, with a re-ranking step
  • Permission to say "I don't know" when the sources don't cover the question
  • Visible citations, so every answer can be checked against its source

Keep reading

Frequently Asked Questions

Is RAG the same as an AI searching the web?

Web search is one possible source. RAG is the general pattern — retrieve, then generate — and it works with any collection of documents: a company wiki, product manuals, support tickets, or a database.

Does RAG stop hallucinations completely?

No. It reduces them by grounding answers in real sources, but if the wrong passage is retrieved or the model misreads it, the answer will still be wrong. Citations make those errors much easier to catch.

Do you need a vector database to build RAG?

Not necessarily. Vector databases are popular for meaning-based search, but small projects can start with an ordinary search index, and many production systems combine keyword and vector search.

TM
Written by

Theo Marsh

Tech & AI Editor

Theo edits the Tech & AI section and checks every piece that goes into it. His background is in fact-checking, which mostly means he's the person who asks "but how do we know that?" until everyone in the room is annoyed. He also writes the internet-literacy guides, and insists the award-show puns were research.