Skip to main content

Context Window

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/context-window

In short

A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.

What is a context window?

The context window is an LLM's working memory for a single request. It covers everything the model can see at once: the system prompt, the conversation so far, any documents or tool results you include, and the answer it is generating. Its size is measured in tokens and varies widely between models, from a few thousand tokens to a million or more.

An LLM has no memory between requests. Chat applications create the feeling of memory by sending the whole conversation again with every new message, so a long chat gradually fills the window. When the limit is reached, the application must drop old messages, summarize them, or return an error; anything that doesn't fit is simply invisible to the model.

An analogy is a desk: a bigger desk lets you spread out more papers at once, but you still can't read the papers left in a filing cabinet in another room. Larger windows let models work with long documents or entire codebases, but each request costs more, responds more slowly, and models can pay less attention to details buried in the middle of a very long context.

A context window is not the model's knowledge. What a model learned during training is stored in its weights, while the context window only holds what you send in the current request. That is why RAG exists: instead of trying to fit every document into the window, it retrieves only the most relevant pieces and places those in the context.

Key takeaways

  • The context window is the maximum number of tokens an LLM can process in one request.
  • It includes the system prompt, history, attached documents, and the model's output.
  • LLMs are stateless, so chat apps resend the conversation with each message.
  • Bigger windows cost more and can make it harder for the model to focus on details.
  • RAG and summarization help fit the most relevant information into the window.

Example

Trimming chat history to fit the context windowtypescript
type Message = { role: string; content: string };
const countTokens = (m: Message) => Math.ceil(m.content.length / 4); // rough estimate

// Keep the system prompt plus as many recent messages as fit in the window
function fitToWindow(messages: Message[], maxTokens: number): Message[] {
  const [system, ...history] = messages;
  let used = countTokens(system);
  const kept: Message[] = [];
  for (const message of history.reverse()) {
    used += countTokens(message);
    if (used > maxTokens) break; // older messages no longer fit and are dropped
    kept.unshift(message);
  }
  return [system, ...kept];
}

Readers ask

What happens when you exceed the context window?

The request either fails with an error or the application has to cut content, usually by dropping or summarizing the oldest messages. The model cannot see anything that was left out, so it may lose track of earlier details.

Is a bigger context window always better?

Not always. A larger window lets you include more material, but each request becomes slower and more expensive, and models can overlook details in very long inputs. Sending only the most relevant information often gives better answers.

Does the context window include the model's answer?

Yes. The input tokens and the generated output tokens share the same window, so a very long prompt leaves less room for the reply. Many APIs also set a separate maximum for output tokens.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings