Context Window
- In Turkish
- Bağlam Penceresi
In short
A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.
What is a context window?
The context window is an LLM's working memory for a single request. It covers everything the model can see at once: the system prompt, the conversation so far, any documents or tool results you include, and the answer it is generating. Its size is measured in tokens and varies widely between models, from a few thousand tokens to a million or more.
An LLM has no memory between requests. Chat applications create the feeling of memory by sending the whole conversation again with every new message, so a long chat gradually fills the window. When the limit is reached, the application must drop old messages, summarize them, or return an error; anything that doesn't fit is simply invisible to the model.
An analogy is a desk: a bigger desk lets you spread out more papers at once, but you still can't read the papers left in a filing cabinet in another room. Larger windows let models work with long documents or entire codebases, but each request costs more, responds more slowly, and models can pay less attention to details buried in the middle of a very long context.
A context window is not the model's knowledge. What a model learned during training is stored in its weights, while the context window only holds what you send in the current request. That is why RAG exists: instead of trying to fit every document into the window, it retrieves only the most relevant pieces and places those in the context.
Key takeaways
- The context window is the maximum number of tokens an LLM can process in one request.
- It includes the system prompt, history, attached documents, and the model's output.
- LLMs are stateless, so chat apps resend the conversation with each message.
- Bigger windows cost more and can make it harder for the model to focus on details.
- RAG and summarization help fit the most relevant information into the window.
Example
type Message = { role: string; content: string };
const countTokens = (m: Message) => Math.ceil(m.content.length / 4); // rough estimate
// Keep the system prompt plus as many recent messages as fit in the window
function fitToWindow(messages: Message[], maxTokens: number): Message[] {
const [system, ...history] = messages;
let used = countTokens(system);
const kept: Message[] = [];
for (const message of history.reverse()) {
used += countTokens(message);
if (used > maxTokens) break; // older messages no longer fit and are dropped
kept.unshift(message);
}
return [system, ...kept];
}Readers ask
What happens when you exceed the context window?
The request either fails with an error or the application has to cut content, usually by dropping or summarizing the oldest messages. The model cannot see anything that was left out, so it may lose track of earlier details.
Is a bigger context window always better?
Not always. A larger window lets you include more material, but each request becomes slower and more expensive, and models can overlook details in very long inputs. Sending only the most relevant information often gives better answers.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- PromptAI & Machine Learning, p. 35A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- Context EngineeringAI & Machine Learning, p. 11Context engineering is the practice of choosing what an LLM sees on each call (instructions, documents, tool results, history) so it can do the task reliably.
- ChunkingAI & Machine Learning, p. 9Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
Spotted a mistake or something missing on this page?Suggest an edit