Context window
The context window is the maximum amount of text (instructions, retrieved information and the conversation so far) that a language model can take into account when producing its next reply. It is measured in tokens.
Also called: context length, token limit
Everything the model "remembers" during a call sits in the context window: the system prompt, any documents or records fetched for the conversation, and the transcript of the exchange so far. What is not in the window does not exist to the model.
A long call, a large knowledge base or a verbose set of instructions all consume the window. When it fills, something has to be dropped or summarized, and that is when an agent may forget a detail given earlier. Well-designed systems manage this deliberately rather than letting the oldest turns fall off the end.
The context window is the mechanism behind turn-based context awareness, the reason a later "the second one" can be resolved against an earlier list. It is a working memory for one conversation, distinct from any memory carried between calls.
