vertex ai memory bank

Vertex AI Memory Bank: Persistent Agent Memory, How It Works, and When to Use It

Vertex AI Memory Bank is the managed long-term memory layer Google Cloud documents for AI agents. This guide maps official docs, the product blog, ADK tools, and a vendor alternative into a checklist and a decision path for when a Google Cloud-bound store fits.

This article was researched with AI assistance and independently reviewed by multiple AI models before publication.

Key takeaways

  • Google Cloud documents Memory Bank as a fully managed, persistent store for personalizing agents and managing the model context window.
  • Google's product blog splits Agent Engine Sessions (history inside one session) from Memory Bank (facts across sessions), and argues that stuffing full transcripts into the context window raises inference cost and latency.
  • ADK exposes MemoryService for persistent long-term memory, InMemoryMemoryService for non-durable local storage, and PreloadMemoryTool plus LoadMemoryTool for retrieval.
  • Virtualization Review reports a July 8, 2025 public-preview announcement, asynchronous Gemini fact extraction, and integration through ADK, LangGraph, or CrewAI.
  • Zep, a vendor, describes Memory Bank as bound to Google Cloud and positions its own product as a multi-LLM, multi-cloud alternative — a residency choice, not a feature bake-off.

Vertex AI Memory Bank is the name Google Cloud's product blog uses for long-term memory that agents can store, retrieve, and manage across multiple sessions. Agent Platform documentation describes Memory Bank as a fully managed, persistent, and accessible memory store that helps you personalize how an agent interacts with users and manage the context window. A Google Developer discussion frames the same capability as long-term memory that sits beyond an LLM's context window.

The rest of this page is a source-scoped briefing: what Memory Bank is documented to do, how it sits next to sessions and ADK tools, and a decision path for whether a Google Cloud-bound memory store matches the job.

What Memory Bank is documented to do

Google announced Memory Bank in public preview on July 8, 2025, as a capability inside Vertex AI Agent Engine, according to Virtualization Review. That write-up describes the design goal as persistent memory for conversational agents — retaining preferences and past choices so interactions can stay personalized across sessions.

The Agent Platform Memory Bank page lists a fully managed, persistent, and accessible store; storage that can be reached from Agent Runtime, a local environment, or other deployment options; and automatic memory revisions so you can inspect how memories change as new information is ingested. It calls long-term personalization the intended fit — for example, a customer service agent that remembers key facts from past support tickets and product preferences instead of asking again.

The developer discussion adds a production framing: Memory Bank is a fully managed service that provides persistent, long-term memory for AI agents. It lists two advantages for builders: you do not have to stand up your own vector databases, retrieval logic, and data pipelines, and it is presented as more efficient at scale than repeatedly stuffing a large context window.

Why not dump the whole transcript into context

Google's public-preview blog states that inserting entire session dialogues into an LLM's context window is expensive and computationally inefficient, with higher inference costs and slower responses. Memory Bank is positioned there as a way to give the agent background on a user without carrying the full transcript.

That same post distinguishes two layers you can enable together:

  • Agent Engine Sessions — store and manage conversation history inside an individual session.
  • Memory Bank — long-term memory so agents can store, retrieve, and manage relevant information across multiple sessions.

A Medium article on Google Cloud makes the same split in ADK terms: long-term memory answers questions such as a favorite programming language or what a customer complained about last time, and is meant to span days, weeks, or months.

How memories are stored and retrieved

The product blog says Memory Bank stores and updates memories intelligently. Key facts — the post's examples are "My preferred temperature is 71 degrees" and "I prefer aisle seats on flights" — are stored persistently and organized by a defined scope such as user ID. Google attributes the learning-and-recall approach to a Google Research method accepted by ACL 2025, described in that post as topic-based.

Virtualization Review reports that Memory Bank processes conversations asynchronously, using Gemini models to extract those kinds of facts, then stores, updates, and retrieves them, including resolving contradictions over time.

The developer discussion argues that true agent memory needs more than storing facts — it also needs the ability to intelligently forget.

ADK, tools, and neighboring frameworks

Google's blog says you can define an agent with the Agent Development Kit (ADK), enable Agent Engine Sessions, then enable Memory Bank.

The Medium ADK write-up describes ADK as a framework for both short-term and long-term memory. Long-term memory is managed through MemoryService, which ingests and retrieves from a persistent knowledge store. ADK also ships InMemoryMemoryService, a lightweight, non-persistent dictionary. Two pre-built retrieval tools are named:

  • PreloadMemoryTool — automatically retrieves relevant memories at the beginning of each conversation turn.
  • LoadMemoryTool — explicit, on-demand retrieval.

Virtualization Review reports that developers can integrate Memory Bank using ADK or through external frameworks such as LangGraph and CrewAI.

A community GitHub plugin documents a separate integration path for agents that call Memory Bank as tools. That README lists prerequisites of a Google Cloud project with billing enabled, the Vertex AI API enabled, an Agent Engine instance created for Memory Bank, and google-cloud-aiplatform at version 1.111.0 or higher. It says install-time config needs projectId, location (example value in the README: us-central1), and reasoningEngineId. After install, the README recommends adding openclaw-vertexai-memorybank to plugins.allow in openclaw.json. The plugin registers four tools the agent can call; the two named in the captured README are memorybank_search (semantic search returning facts with similarity scores, topics, timestamps, and memory IDs) and memorybank_forget (delete a memory by ID). Introspection can include similarity scores by default, or return facts only. The same README describes a free tier of 1,000 retrievals per month; that figure is from the plugin page, not from a captured Google Cloud pricing document.

A Google Colab notebook is linked from Google's Memory Bank materials. The captured extract from that notebook is only the title fragment "Google Colab"; the cited material does not describe notebook contents or a getting-started path.

Where Memory Bank sits on Agent Platform

Memory Bank is one piece of a broader production story. Gemini Enterprise Agent Platform scale docs describe a fully managed environment for testing, release management, and reliability, plus Example Store and Evaluation Service for a feedback loop, Cloud Trace for response times and executed operations, and IAM agent identity for security and access management.

Source-scoped checklist

QuestionWhat the cited sources say
What problem is named?Persistent memory across sessions; avoid stuffing full transcripts into the context window (Google blog, Virtualization Review, developer discussion).
What is stored?Extracted facts scoped by identifiers such as user ID; examples include temperature and seat preferences (Google blog, Virtualization Review).
How is it managed?Fully managed store with revisions; ADK MemoryService plus preload/on-demand tools (Agent Platform docs, Medium ADK article).
What is in-session vs long-term?Agent Engine Sessions for history inside a session; Memory Bank for facts across sessions (Google blog).
What residency constraint is claimed?Zep, a vendor, states Memory Bank is bound to Google Cloud.
What official Memory Bank price is in this evidence set?This evidence bundle did not include a captured Google Cloud pricing page for Memory Bank.

Those rows are source-scoped summaries, not a field-wide consensus. Figures that look conflicting in a research dump (research method vs preview date vs SDK install) sit at different layers and can coexist; they are not treated here as disagreements.

Decision path: does Memory Bank fit this job?

Walk the tree in order. Stop at the first match.

  1. You only need recall inside one live conversation. Enable session history (Agent Engine Sessions in Google's blog; ADK short-term memory in the Medium article). Memory Bank is the wrong layer.
  2. You need facts that survive across days or months — preferences, prior tickets — and you already run agents on Google Cloud. Memory Bank is the layer Google documents for that: managed persistence, scoped facts, revisions, and ADK or LangGraph/CrewAI integration.
  3. You want a managed memory service so you are not operating vector stores and retrieval pipelines yourself, and Google Cloud residency is acceptable. The developer discussion's "undifferentiated work" argument applies. Confirm IAM agent identity and tracing on Agent Platform before production traffic.
  4. You need the memory plane to stay multi-cloud or multi-LLM rather than Google-bound. Zep's comparison page frames that as lock-in versus neutrality. Treat it as a residency decision, not as an independent bake-off of retrieval quality.
  5. You are wiring an existing agent runtime that already speaks tools, and you can create an Agent Engine instance. The community plugin README is one documented adapter (search, forget, score metadata). Verify the SDK pin and the three config fields against your project.

Adjacent industry notes (not Memory Bank itself)

Virtualization Review places the preview in a wider shift: it highlights Claude Opus 4 as skilled at creating and maintaining "memory files," and notes Microsoft announced a Copilot feature named Memory for user-specific information across sessions. Those are neighboring products, not Vertex AI Memory Bank capabilities.

Sources