Chunking
- Pronunciation
- CHUNK-ing
In short
Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
What is chunking in RAG?
A retrieval-augmented generation (RAG) system can't hand a whole manual to a language model, and a single embedding of a long document blurs many topics into one vector. So documents are first split into chunks, usually a few hundred tokens each. Every chunk gets its own embedding and is stored in a vector database, and at question time the system retrieves the few chunks closest to the question and adds them to the prompt.
How you split matters. Fixed-size chunking cuts the text every so many tokens, often with an overlap of 10 to 20 percent, so a sentence cut in half still appears whole in one chunk. Structure-aware chunking prefers natural boundaries, such as headings, paragraphs and sentences, or functions in source code. The early research set a simple baseline: Dense Passage Retrieval and the original RAG paper both split Wikipedia into passages of 100 words.
Chunk size is a trade-off. Small chunks match a question precisely but can lose the context that explains them; large ones keep the context but dilute the match and fill the context window faster. A common fix is to store each chunk with metadata, such as its title, section and source address, or to retrieve small chunks and pass the larger section around them to the model. Trying a few sizes against real questions is the usual way to choose.
Key takeaways
- Chunking splits documents into passages that are embedded and retrieved one by one.
- Chunks are usually a few hundred tokens long, often with some overlap.
- Splitting at headings, paragraphs or sentences keeps ideas whole.
- Smaller chunks match precisely; larger ones keep more context.
Example
def chunk(words, size=200, overlap=40):
"""Split a list of words into chunks of `size` words that share `overlap` words."""
step = size - overlap
return [words[i:i + size] for i in range(0, max(len(words) - overlap, 1), step)]
words = ("lorem " * 1000).split() # 1,000 words
chunks = chunk(words)
print(len(chunks)) # 6
print([len(c) for c in chunks]) # [200, 200, 200, 200, 200, 200]Readers ask
What chunk size should I use?
There is no single right answer; it depends on the documents and the questions. A few hundred tokens with a small overlap is a common starting point. Try two or three sizes on a set of real questions and keep the one whose retrieved chunks answer them best.
Why not put the whole document in the prompt?
Long prompts cost more, run slower and can exceed the context window, and models tend to use information at the start and end of a long context better than in the middle. Retrieving only the relevant chunks keeps the prompt short and focused.
See also
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- Context WindowAI & Machine Learning, p. 12A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.
- TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- Semantic SearchAI & Machine Learning, p. 42Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
Sources
- Karpukhin et al.: Dense Passage Retrieval for Open-Domain Question Answering (2020)arxiv.org(opens in a new tab)
- Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020)arxiv.org(opens in a new tab)
- Liu et al.: Lost in the Middle: How Language Models Use Long Contexts (2023)arxiv.org(opens in a new tab)
Spotted a mistake or something missing on this page?Suggest an edit