A context window is the bounded sequence of information a language model can use when generating an output. In large language models, it is usually measured in tokens and contains the current input, relevant conversation history, and, under common accounting conventions, the output being generated. It is often described as working memory, but it is distinct from knowledge encoded in model parameters through training data. A larger window permits more material to be supplied together; it does not guarantee that every detail will be used accurately. (platform.claude.com)
Measurement and contents
Tokens are the units produced by tokenization. Depending on the tokenizer, they may represent words, word fragments, punctuation, or byte-based units. Consequently, context length is not equivalent to a fixed number of words, characters, or pages. The same passage can occupy different numbers of tokens under different tokenizers, and languages and text formats differ in their token requirements. (huggingface.co)
In conversational systems, context may include system instructions, user messages, previous assistant responses, tool definitions, and tool results. Systems supporting multimodal learning may also account for images and other nontextual inputs. The visible chat transcript therefore need not represent the entire context supplied to the model. Exact accounting depends on the model and interface. (platform.claude.com)
Where input and output share a total limit , the basic constraint is:
Thus, a hypothetical 32,000-token window containing 28,000 input tokens leaves at most 4,000 tokens for generation, before any additional interface-specific constraints. Context capacity and maximum output length are separate specifications: a large input allowance does not necessarily permit an equally long answer. (platform.claude.com)
Architectural basis
The Transformer architecture, introduced in 2017, processes sequences using an attention mechanism. In self-attention, token representations are combined with information from other positions. A causal decoder masks future positions, allowing each generated token to depend on preceding tokens without accessing later ones. Positional encoding supplies information about sequence order that attention alone does not explicitly represent. (arxiv.org)
Long sequences create substantial computational costs. For fixed representation dimensions, conventional dense attention performs work proportional to the square of sequence length; a straightforward implementation also materializes a quadratically sized attention matrix. FlashAttention reorganizes exact attention computation to reduce transfers between levels of GPU memory and avoid storing the full attention matrix. It improves practical efficiency without making dense attention’s arithmetic inherently linear in sequence length. (arxiv.org)
During autoregressive inference, a key–value cache stores previously computed attention keys and values. Reusing these representations avoids repeatedly processing earlier tokens, but the cache consumes memory. In full-attention layers, cache requirements grow with the retained sequence; sliding-window attention limits direct attention to a bounded region. A local attention window is therefore not necessarily identical to the total sequence length an architecture can process. (huggingface.co)
Extending context length
A usable context limit depends on more than available memory. Positional representations and the sequence lengths encountered during training also matter. Extending a configuration’s nominal maximum does not by itself establish reliable performance beyond the lengths for which a model was prepared. Methods such as YaRN modify rotary positional representations and use additional training to extend supported context lengths efficiently. (arxiv.org)
Context extension is distinct from increasing parameter count. It changes how much information can be supplied in a sequence rather than directly specifying the amount of knowledge stored in model weights. Similarly, in-context learning uses instructions or demonstrations within the input without updating parameters, whereas fine-tuning changes model parameters through training. A larger window can accommodate more demonstrations, but their presence alone does not ensure improved task performance. (platform.claude.com)
Nominal capacity and effective use
A nominal context window states how much input a system supports, not how reliably it reasons over that input. The 2023 study Lost in the Middle examined multi-document question answering and key–value retrieval. In the evaluated models, performance often improved when relevant information appeared near the beginning or end and declined when it appeared in the middle. These findings describe particular experiments rather than a universal rule for all models. (arxiv.org)
Long-context evaluation therefore distinguishes acceptance of a sequence from successful information retrieval and reasoning within it. A “needle-in-a-haystack” test checks whether a model can locate a small piece of information among distractors. The RULER benchmark broadens this approach with multiple targets, multi-hop tracing, and aggregation tasks. Its experiments showed that strong performance on simple retrieval could coexist with substantial deterioration as context length or task complexity increased. Effective context length is consequently task- and evaluation-dependent. (arxiv.org)
Context management and external memory
Applications can manage growing conversations by retaining selected messages, removing older material, or replacing portions with summaries. Such compaction changes the information available for subsequent generation; it does not enlarge the model’s underlying window. A conversation stored by an application can therefore be much longer than the material included in any individual request. (platform.claude.com)
Retrieval-augmented generation provides a different form of access to information. A retriever selects relevant passages from an external knowledge base, and the generator conditions its answer on those passages. The collection can greatly exceed the context window because only selected material is supplied at a time. This separates persistent external storage from temporary model context, while leaving retrieval quality and the model’s use of retrieved evidence as distinct constraints. (arxiv.org)
References
- Context windows - Claude Platform Docsplatform.claude.com
- Tokenization algorithms · Hugging Facehuggingface.co
- Attention Is All You Needarxiv.org
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- Cache strategies · Hugging Facehuggingface.co
- YaRN: Efficient Context Window Extension of Large Language Modelsarxiv.org
- Language Models are Few-Shot Learnersarxiv.org
- Lost in the Middle: How Language Models Use Long Contextsarxiv.org
- RULER: What's the Real Context Size of Your Long-Context Language Models?arxiv.org
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasksarxiv.org