Skip to content
How It Thinks

Design preview · Sample content — not a released episode or final script.

Memory & context

Context and memory solve different problems

What a model can see during a response is different from what a system stores and retrieves between responses.

Sample reading time: 2 minutes

Conceptual illustration: a row of stored cards, one tagged orange, feeds a single item into a green frame of current context.
Conceptual illustration

Summary

Context is the information presented for a model invocation. Memory is a system-level method for retaining and selecting information across time. More stored information does not automatically mean more useful context.

Video embed placeholder — no video attached

An assistant can appear to remember something because it is present in the current conversation, because software retrieved it from storage, or because the fact is part of the model's learned parameters. Those routes can produce similar-looking answers while relying on different mechanisms. Separating them makes it easier to reason about reliability and control.

Context is the working input

A model invocation receives an input that can include instructions, conversation messages, documents and tool results. That collection is its context for the response. The input has practical size limits, so a system must decide what to include. Putting every available document into one prompt is not always feasible or useful. Relevant details can be obscured by material that does not help answer the question.

Storage is not retrieval

A memory system can retain notes, summaries or records outside the current prompt. To affect a later response, some of that information must be selected and brought back into context. The retrieval step is therefore as important as the storage step. A system might search by meaning, use identifiers, apply filters or combine several methods. Each choice changes what the model sees. A stored record that is never retrieved cannot help the current answer.

A simplified mechanism

  1. Stored records
  2. Retrieve relevant itemsConversation input joins Assemble current context
  3. Assemble current context
  4. Model response
Diagram description: Stored records flow through retrieval into the current context and then to a model response. A separate conversation input also joins the current context.

Summaries trade detail for space

Summarizing earlier material can reduce the amount of text a model needs to read. It also discards detail. A summary may preserve the conclusion but lose the exception, the source or the reason a decision was made. For tasks where exact wording matters, the system may need to return to the original record. Good memory design makes it possible to inspect where a remembered claim came from and whether it is still current.

The useful questions to ask

What gets stored? How is it selected later? Can a person inspect or correct it? Does the system distinguish an instruction from quoted material? These questions are more informative than asking whether the assistant has memory at all. In this sample architecture, memory is an information pipeline around a model, not a promise of perfect recall. Understanding that pipeline helps explain why an assistant can remember one detail and miss another.

Design preview

← Back to all topics