RAG
Retrieval-Augmented Generation
- Pronunciation
- RAG
In short
RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
What is RAG?
Retrieval-Augmented Generation, or RAG, combines a search step with a language model. When a user asks a question, the system first retrieves the most relevant pieces of information from a knowledge source, such as company documents or a help center, and adds them to the prompt. The LLM then generates its answer from that supplied context instead of relying only on what it learned during training.
A typical RAG pipeline has two phases. Ahead of time, documents are split into smaller chunks, each chunk is turned into an embedding, and the embeddings are stored in a vector index. At question time, the question is embedded too, the closest chunks are retrieved, and they are inserted into the prompt together with an instruction such as 'answer using only the sources below'.
An everyday analogy is an open-book exam: instead of answering from memory, the student looks up the right pages first. RAG is widely used for internal knowledge assistants, customer support bots, and documentation search, because it reduces hallucinations and lets answers point to their sources.
RAG is often compared with fine-tuning. Fine-tuning trains the model further on your data, which changes its behavior or style but is slow and costly to update, while RAG leaves the model unchanged and simply gives it fresh information on every request. For knowledge that changes often, RAG is usually the simpler and cheaper choice.
At a glance
Key takeaways
- RAG retrieves relevant data first, then asks the LLM to answer using it.
- It usually relies on embeddings and a vector index to find relevant text chunks.
- It keeps answers current without retraining the model.
- It reduces, but does not eliminate, hallucinations.
- The quality of retrieval largely determines the quality of the answer.
Example
// embed, vectorIndex and llm are placeholders for real services
async function answer(question: string): Promise<string> {
// 1. Retrieve: find the text chunks most similar to the question
const queryVector = await embed(question);
const chunks = await vectorIndex.search(queryVector, { topK: 3 });
// 2. Augment: add the retrieved text to the prompt
const sources = chunks.map((chunk) => chunk.text).join("\n\n");
const prompt = `Answer using only these sources:\n${sources}\n\nQuestion: ${question}`;
// 3. Generate: the LLM writes an answer grounded in the sources
return llm.generate(prompt);
}Readers ask
What is the difference between RAG and fine-tuning?
RAG gives the model relevant information at request time without changing the model, while fine-tuning trains the model further on your own data. RAG is better for facts that change often; fine-tuning is better for teaching a consistent style, format, or specialized behavior.
Does RAG stop hallucinations?
RAG reduces hallucinations by giving the model real sources to work from, but it does not eliminate them. The model can still misread the sources, and if retrieval returns the wrong documents, the answer can still be wrong.
Do I need a vector database for RAG?
Not necessarily. Many RAG systems use a vector database, but a regular database with vector search support, or even classic keyword search, can also serve as the retrieval step.
Often compared
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- PromptAI & Machine Learning, p. 35A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- HallucinationAI & Machine Learning, p. 23A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
- Database IndexDatabases, p. 7A database index is a data structure that helps a database find rows quickly without scanning a whole table, much like the index at the back of a book.
- Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- Cosine SimilarityAI & Machine Learning, p. 13Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
- EvalsAI & Machine Learning, p. 17Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
- ChunkingAI & Machine Learning, p. 9Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
Sources
Spotted a mistake or something missing on this page?Suggest an edit