memory in the age of ai agents

Memory in the Age of AI Agents: What the Survey and the Explainers Actually Say

A survey released in December 2025, a maintained companion paper list, and a set of published explainers now describe agent memory in public. This guide synthesizes what those sources say about definitions, memory types, memory operations, enterprise constraints, and where their vocabulary diverges — attributing each point to the source that makes it.

This article was researched with AI assistance and independently reviewed by multiple AI models before publication.

Key takeaways

  • Kore.ai's explainer describes a model without memory as something that "processes what is placed in front of it, generates a response, and retains nothing," and defines agent memory as the system that encodes, stores durably, retrieves, and updates information.
  • IBM's explainer defines AI agent memory as the ability to store and recall past experiences to improve decision-making, perception and overall performance, in contrast to models that process each task independently.
  • Atlan's overview names four memory types — in-context (working), episodic, semantic, and procedural — and attributes their formalisation for LLMs to the CoALA framework from Princeton (arXiv:2309.02427).
  • IBM and Atlan both trace their taxonomy to CoALA but label the immediate layer differently: "short-term memory" versus "in-context (working) memory."
  • An arXiv survey positions itself as an up-to-date landscape of agent memory research and compiles a summary of memory benchmarks and open-source frameworks; its companion GitHub paper list records the survey's release on 2025/12/16 and reported reaching 1k stars in a 2026/01/29 note.

A model with no memory, in the words of Kore.ai's explainer on agent memory, "processes what is placed in front of it, generates a response, and retains nothing." That same page opens with a blunt framing: "Agent memory is not a feature." This guide treats agent memory as the machinery that changes that baseline, and every point below is attributed to the published source that states it.

What follows is a synthesis of published material — a research survey, its companion paper list, and several explainers. Attribution matters here, because these sources use different vocabulary and, in one case, propose a different category count.

What "agent memory" is being defined as

Two definitions from the gathered material are worth putting side by side.

Kore.ai defines AI agent memory operationally: it is the system that allows an agent to encode information from interactions, store it durably, retrieve it when relevant, and update it as conditions change. That is four verbs, and they map neatly onto four places an implementation can break.

IBM's Think explainer, written by Cole Stryker, defines it by outcome: the ability of an AI system to store and recall past experiences to improve decision-making, perception and overall performance. IBM contrasts this with traditional AI models that process each task independently, noting that agents with memory can retain context, recognize patterns over time and adapt based on past interactions. Its illustration is a system that, instead of reacting only to the current temperature, stores and analyzes past data to make more intelligent decisions.

Read together, one definition tells you what the machinery does and the other tells you what it buys you.

The taxonomy IBM and Atlan both borrow

Two of the explainers gathered here trace their category scheme back to a single academic source.

Atlan's overview of agent memory types names four: in-context (working) memory, episodic memory, semantic memory, and procedural memory — drawn from cognitive science and formalised for LLMs in the CoALA framework (Princeton, arXiv:2309.02427). Each stores a different class of information: the live context window, past events, factual knowledge, and behavioural rules respectively. Atlan dates the CoALA paper to 2023 and credits Endel Tulving's 1972 distinction between episodic and semantic memory with giving AI researchers a ready-made framework.

IBM attributes its own breakdown to the same Cognitive Architectures for Language Agents (CoALA) paper from a team at Princeton University, and begins with short-term memory (STM) — which, in IBM's description, enables an agent to remember recent inputs for immediate decision-making. Its example: a chatbot that remembers previous messages within a session can give coherent responses instead of treating each user input in isolation.

So the immediate layer carries two names across these two explainers — IBM's "short-term memory" and Atlan's "in-context (working) memory" — while both cite CoALA as the origin. If you are comparing write-ups, expect vocabulary drift before you assume you are looking at competing models.

Where one source extends the standard list

Atlan argues the four-type taxonomy is incomplete for one specific setting: enterprise data agents running against live data estates require a fifth type the standard taxonomy omits, which it calls organisational context memory. Atlan points to arXiv:2603.17787 (Governed Memory, March 2026) as support for that fifth type.

This is one publisher's proposed extension for one deployment context, not a settled category. Treat it as a question to ask about your own system — does the agent need to know things about the organisation that live outside any conversation? — rather than as a fifth box everyone has agreed to.

Why the topic surfaced when it did

The phrase "memory in the age of AI agents" is itself the title of a research survey. The arXiv paper states that the work "aims to provide an up-to-date landscape of current agent memory research" and that, to support practical development, the authors "compile a comprehensive summary of memory benchmarks and open-source frameworks." The same abstract appears on its Hugging Face paper page.

The survey's companion GitHub paper list records the timeline in its own news log: the repository and the survey both went up on 2025/12/16, with the paper featured that day as Hugging Face Daily Paper #1; a 2026/01/13 entry notes the survey was updated to incorporate several recent works; and a 2026/01/29 entry announces the repository reached 1k stars.

That compiled benchmark-and-framework summary is the most useful artifact in this list for anyone building. It is the difference between reading about memory types and having a list of things to evaluate against.

The agreement checklist

Across the material gathered here, a few points recur, each traceable to a specific publisher:

  • Statelessness is the starting condition. Kore.ai describes the memoryless baseline as retaining nothing between turns.
  • Memory is a multi-step system, not a store. Kore.ai breaks it into encode, store durably, retrieve when relevant, update as conditions change.
  • Both explainers trace the category scheme to CoALA. IBM and Atlan each attribute their memory types to that Princeton paper.
  • The scheme is borrowed from cognitive science. Atlan traces the episodic/semantic split to Tulving's 1972 distinction.
  • Selectivity is part of the design, not an optimisation afterthought. IBM states that optimized memory management helps ensure systems store only the most relevant information while maintaining low-latency processing for real-time applications.
  • The research side has consolidated. The arXiv survey presents itself as a landscape review with a compiled summary of benchmarks and open-source frameworks.

A reading path from your situation to a next step

The cited material describes memory types and operations rather than ranking them for particular projects, so the mapping below is inference drawn from those descriptions — a way to route yourself, not a finding any of these sources states.

  • Your agent loses the thread inside a single conversation. That is the layer IBM illustrates with a chatbot that remembers previous messages within a session — its short-term memory, Atlan's in-context (working) memory. Start there before adding storage.
  • Your agent forgets specific past events across sessions. Atlan assigns past events to episodic memory and factual knowledge to semantic memory — two different stores with two different retrieval questions.
  • Your agent keeps relearning how to do a task. Atlan places behavioural rules in procedural memory.
  • Your agent works against live enterprise data estates. This is exactly the gap Atlan says the standard four-type taxonomy omits, and where it proposes organisational context memory.
  • You need to evaluate implementations rather than concepts. Go to the compiled benchmark and open-source framework summary in the arXiv survey and the maintained paper list.
  • You are diagnosing why a pilot did not become a deployment. Kore.ai offers one candidate explanation, discussed next.

The enterprise framing

Some of this material puts memory in an organisational rather than purely architectural frame.

Kore.ai states that context discontinuity — the failure of AI agents to retain user context across interactions — is consistently cited as one of the primary reasons AI deployments stall after initial pilots. That is Kore.ai's characterization of a commonly cited reason; the captured material presents it as a widely made observation rather than as a measured result, and it is worth reading as a hypothesis to check against your own deployment data.

Dataiku's overview is titled "AI agent memory: types, architectures, and enterprise considerations," published August 18, 2026 and credited to Team Dataiku. The title names three things — types, architectures, and enterprise considerations — which is the extent of what the material cited here establishes about it; the piece is listed as a further reading pointer, not summarised.

On the governance edge, Atlan cites arXiv:2603.17787 (Governed Memory, March 2026) alongside its organisational context memory proposal — an indication that who is allowed to remember what has started to be treated as its own research question, not only an engineering one.

What this guide deliberately leaves out

The sources synthesized here define memory, categorise it, and point to benchmarks and frameworks. Concrete implementation choices — specific storage engines, retrieval tuning, cost curves, or measured performance comparisons — are not established by the material cited above, so this guide does not assert them. For that layer, the compiled benchmark and framework summary in the arXiv survey is the appropriate next stop.

Sources