AI Agents for Beginners - 13. Memory for AI Agents
Memory types for AI Agents (Working, Short-term, Long-term, Persona, Episodic, Entity, Structured RAG), implementation with Mem0, Cognee, and Azure AI Search, and the self-improving agent pattern with code examples.
June 9, 2026
AI Agents for Beginners - 13. Memory for AI Agents
This article summarizes Lesson 13 of Microsoft's AI Agents for Beginners course.
Two unique advantages frequently highlighted for AI Agents are their ability to complete tasks by invoking tools and their ability to improve over time.
Memory is the foundation for building self-improving agents that deliver better user experiences.
Memory is the foundation for building self-improving agents that deliver better user experiences.
Understanding AI Agent Memory
At its core, AI Agent memory is the mechanism that enables retaining and recalling information.
This information can include conversation details, user preferences, past actions, or learned patterns.
This information can include conversation details, user preferences, past actions, or learned patterns.
Without memory, AI applications are often stateless, starting every interaction from scratch. This leads to a repetitive and frustrating user experience where the agent "forgets" previous context or preferences.
Why is Memory Important?
An agent's intelligence is deeply tied to its ability to recall and utilize past information. Memory enables agents to become:
| Characteristic | Description |
|---|---|
| Reflective | Learns from past actions and outcomes |
| Interactive | Maintains context across continuous conversations |
| Proactive and Reactive | Anticipates needs or responds appropriately based on historical data |
| Autonomous | Operates more independently using stored knowledge |
The goal of implementing memory is to make agents more reliable and capable.
Types of Memory
| Type | Lifespan | Description | Travel Example |
|---|---|---|---|
| Working Memory | Single task | Immediate information needed for the next step. Captures requirements, decisions, and actions | "I want to book a trip to Paris" — Holds current request in immediate context |
| Short-term Memory | Single session | Current conversation context. Can reference previous turns | "Flights to Paris?" → "What about lodging there?" — Remembers that "there" is Paris |
| Long-term Memory | Persists across sessions | User preferences, history, and general knowledge. The core of personalization | "Ben likes skiing, prefers coffee with a mountain view, avoids advanced slopes due to injury" |
| Persona Memory | Persistent | Maintains a consistent role and persona for the agent | Reinforces "expert ski planner" role — Tailors responses with expert tone and domain knowledge |
| Episodic Memory | Persistent | Step-by-step success and failure sequences of tasks. Experience-based learning | Records a booking failure on a specific flight (no seats) → Automatically attempts alternative flights |
| Entity Memory | Persistent | Extracts and structures people, places, and things from conversations | Extracts "Paris", "Eiffel Tower", "Le Chat Noir" → Suggests rebooking in the future |
| Structured RAG | Persistent | Extracts structured information from diverse sources. Supports precise queries | Parses flight details from email → Enables answering "When was the Tuesday flight to Paris booked?" |
Process FlowchartMermaid%%{init: {'look': 'handDrawn', 'flowchart': {'subGraphTitleMargin': {'top': 12, 'bottom': 6}}, 'themeVariables': {'clusterBkg': '#e8efff28', 'clusterBorder': '#aabbcc'}}}%% flowchart LR Root["Memory Types"] Root --> WM["Working\nMemory"] Root --> ST["Short-term\nMemory"] Root --> LT["Long-term\nMemory"] Root --> PM["Persona\nMemory"] Root --> EM["Episodic\nMemory"] Root --> ENT["Entity\nMemory"] Root --> SR["Structured\nRAG"]
Implementing and Storing Memory
Implementing AI Agent memory involves a systematic process of memory management — including creation, storage, retrieval, consolidation, updating, and even "forgetting" (deletion).
Specialized Memory Tools
Mem0
One way to store and manage agent memory is to use specialized tools like Mem0. Mem0 acts as a persistent memory layer, allowing agents to recall relevant interactions, store user preferences and factual context, and learn from successes and failures over time.
Mem0 operates with a two-stage memory pipeline: extraction and update. First, messages added to the agent thread are sent to the Mem0 service, where an LLM summarizes conversation history to extract new memories. Afterwards, an LLM-based update step decides whether to add, modify, or delete these memories, storing them in a hybrid datastore encompassing vector, graph, and key-value databases.
Mem0: Two-Stage Memory PipelineMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR Msg["Agent Messages"] --> Mem0["Mem0 Service"] Mem0 --> LLM["LLM\nSummarize & Extract Conversation"] LLM --> Dec{"Add / Modify\n/ Delete?"} Dec --> Store["Hybrid Datastore\nVector + Graph + Key-Value"] Store -->|"recall"| Agent["Agent"]
Cognee
Another powerful approach is Cognee — an open-source semantic memory that transforms structured and unstructured data into an embedding-based, queryable Knowledge Graph.
Cognee provides a dual-store architecture combining vector similarity search and graph relationships, enabling the agent to understand not only which information is similar, but also how concepts relate to one another.
| Feature | Description |
|---|---|
| Hybrid Retrieval | Blends vector similarity, graph structure, and LLM reasoning — from chunk lookup to graph-aware QA |
| Living Memory | Evolves and grows while remaining queryable as a single interconnected graph |
| Dual-store | Supports both short-term session context and long-term persistent memory |
Cognee: Dual-Store ArchitectureMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR Data["Structured /\nUnstructured Data"] Data -->|"cognee.add()"| VEC["Vector Store\nSimilarity Search"] Data -->|"cognee.cognify()"| GR["Graph Store\nRelations & Structure"] VEC & GR --> HQ["Hybrid Query\n(GRAPH_COMPLETION)"] HQ -->|"relevant context"| Agent["Agent"]
Storing Memory with RAG
Beyond specialized memory tools, Azure AI Search can be utilized as a memory storage and retrieval backend — particularly suited for Structured RAG.
Azure AI Search supports Structured RAG to extract and retrieve dense, structured information from large-scale datasets such as conversation history, emails, and images. It delivers "superhuman precision and recall" compared to traditional text chunking and embedding approaches.
Storing Memory with RAG (Azure AI Search)Mermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR D1["Conversation History"] & D2["Emails"] & D3["Images"] --> AIS["Azure AI Search"] AIS -->|"Structure & Index"| IDX["Structured Index\nDense Information Extraction"] IDX -->|"Precise Query"| Agent["Agent\nSuperhuman Precision & Recall"]
Working Memory with Sessions
setup.pypython
Working memory via sessions — by passing the same session to
agent.run() calls, the agent can see the full conversation history:working_memory.pypython
Creating a new session causes the agent to forget the previous conversation:
new_session.pypython
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR U1["User: 'Love beaches, budget $3000'"] -->|"session"| A["TravelMemoryAgent"] U2["User: 'What did I say my budget was?'"] -->|"same session"| A A -->|"Reference previous turn"| R["Accurately recalls $3000"] U3["User: 'What is my budget?'"] -->|"new_session"| B["TravelMemoryAgent\n(New session)"] B -->|"No previous conversation"| NR["Unknown — New session"]
Long-Term Memory Pattern
To remember user preferences across sessions, a persistent store outside the conversation thread is required. The agent accesses this store via tools:
memory_tools.pypython
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR subgraph Agent["MAF Agent (LLM)"] direction TB LLM["LLM"] end subgraph Tools["@tool functions"] direction TB SP["save_preference()"] GP["get_preferences()"] SH["search_hotels()"] end subgraph Store["Preference Store (dict)"] direction TB PS["user_id → [prefs]"] end LLM -->|"call"| Tools SP & GP --> PS PS -->|"persists across\nsessions"| PS SH -->|"hotel DB"| LLM
Scenario 1 — First-time user books an anniversary trip
Sarah visits for the first time. The agent stores preferences and recommends hotels using tools:
scenario1_session1.pypython
scenario1_session1_followup.pypython
Scenario 2 — Sarah returns weeks later
Sarah returns in a new thread (new session). Working memory is empty, but information remains in the long-term preference store:
scenario2_session2.pypython
scenario2_followup.pypython
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD subgraph S1["Session 1 — First Visit"] U1["Sarah: 'Romantic, spa, accessible\nBudget $700-800/night'"] U2["Sarah: 'Vegetarian, nut allergy'"] U1 & U2 -->|"save_preference()"| STORE["preference_store\n['romantic', 'spa', 'accessible',\n 'vegetarian', 'nut allergy']"] end subgraph S2["Session 2 — Returns Weeks Later"] U3["Sarah: 'Recommend a good hotel'\n(New session, no working memory)"] U3 -->|"get_preferences()"| STORE STORE -->|"Return saved preferences"| A2["TravelBookingAssistant\nPersonalized recommendations"] end
Storing Memory with Cognee Knowledge Graphs
Cognee transforms unstructured text into a queryable knowledge graph, providing agents with relation-aware long-term memory.
Part 1 — Building the Knowledge Base
We collect three types of data to build a comprehensive knowledge base:
cognee_setup.pypython
cognee_data.pypython
cognee_ingest.pypython
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD D1["Developer Profile\n(developer_intro)"] D2["Past Conversations\n(human_agent_conversations)"] D3["Python Principles\n(python_zen_principles)"] D1 & D2 -->|"cognee.add(node_set=['developer_data'])"| ADD["cognee.add()"] D3 -->|"cognee.add(node_set=['principles_data'])"| ADD ADD -->|"cognee.cognify()"| KG["Knowledge Graph\n+ Vector Embeddings"] KG -->|"cognee.memify()"| Rules["Enriched Memory\nPatterns, Rules & Relations"] KG -->|"visualize_graph()"| HTML["cognee_graph.html\nInteractive Visualization"]
Part 2 — MAF Agent with Cognee Tools
The MAF agent queries the Cognee knowledge graph via
@tool functions:cognee_tools.pypython
cognee_agent.pypython
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR subgraph MAF["MAF Agent (AgentSession)"] direction TB LLM2["CodingAssistant\n(LLM)"] Sess["create_session()\nWorking Memory"] LLM2 <--> Sess end subgraph CogneeTools["@tool functions"] direction TB SK["search_knowledge()\nGRAPH_COMPLETION"] SP2["search_principles()\nprinciples_data only"] end subgraph KG2["Cognee Knowledge Graph"] direction TB VEC["Vector Embeddings"] GR["Graph Relations"] end LLM2 -->|"call"| CogneeTools CogneeTools -->|"cognee.search()"| KG2 KG2 -->|"Return relevant knowledge"| LLM2
Making AI Agents Self-Improve
A common pattern for self-improving agents is introducing a "knowledge agent". This separate agent observes conversations between the user and the primary agent, performing the following roles:
- Identify valuable information: Determines which parts of the conversation are worth storing as general knowledge or specific user preferences
- Extract and summarize: Extracts and summarizes essential learnings or preferences from the conversation
- Store in a knowledge base: Stores the extracted information in a vector database or similar storage
- Augment future queries: When the user initiates a new query, retrieves relevant stored information and appends it to the user prompt (similar to RAG)
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD User["User"] -->|"New query"| KA["Knowledge Agent\n(Observer)"] KA -->|"Retrieve relevant knowledge\n(RAG)"| VDB["Vector DB\nKnowledge Base"] VDB -->|"Add stored context"| PA["Primary Agent"] PA -->|"Response"| User PA -->|"Observe conversation"| KA KA -->|"Valuable info\nExtract, summarize & store"| VDB
Optimizations for Memory
| Optimization | Description |
|---|---|
| Latency Management | Quickly verify the value of storing or retrieving information using cheaper, faster models first, calling complex extraction and retrieval processes only when needed |
| Knowledge Base Maintenance | Manage costs by moving less frequently used information to "cold storage" as the knowledge base grows |
Summary
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR Root["Memory for\nAI Agents"] Root --> Types["Memory Types"] Root --> Tools["Memory Tools"] Root --> Pattern["Self-Improving\nPattern"] Types --> T1["Working\ncreate_session()"] Types --> T2["Short-term\nIn-session context"] Types --> T3["Long-term\n@tool persistent store"] Types --> T4["Persona·Episodic\nEntity·Structured RAG"] Tools --> TL1["Mem0\nExtract + Update Pipeline"] Tools --> TL2["Cognee\nKnowledge Graph + Vector"] Tools --> TL3["Azure AI Search\nStructured RAG Backend"] Pattern --> P1["Knowledge Agent\nObserve & Extract Conversation"] Pattern --> P2["RAG Augmentation\nAdd context to future queries"]
| Memory Type | MAF Mechanism | Lifespan |
|---|---|---|
| Working | agent.create_session() | Single conversation |
| Short-term | Cumulative context within thread | Single task/session |
| Long-term | External store accessed via @tool functions | Persists across sessions |
- Provides working memory via
agent.create_session()— the agent sees the full conversation history within the session - New sessions lose context — without long-term memory, the agent cannot recall past conversations
@toolfunctions bridge the gap — allowing the agent to store and retrieve information from a persistent store- Mem0: Transitions stateless agents into stateful ones using a two-stage pipeline (extraction → update)
- Cognee: Relation-aware long-term memory via knowledge graphs + vector embeddings — understands relationships between concepts using
GRAPH_COMPLETIONsearch - Forms a self-improving loop where the more preferences and history are stored, the more refined the agent's recommendations become