AI Agents for Beginners - 13. Memory for AI Agents

Memory types for AI Agents (Working, Short-term, Long-term, Persona, Episodic, Entity, Structured RAG), implementation with Mem0, Cognee, and Azure AI Search, and the self-improving agent pattern with code examples.
June 9, 2026

AI Agents for Beginners - 13. Memory for AI Agents

This article summarizes Lesson 13 of Microsoft's AI Agents for Beginners course.
Two unique advantages frequently highlighted for AI Agents are their ability to complete tasks by invoking tools and their ability to improve over time.
Memory is the foundation for building self-improving agents that deliver better user experiences.

Understanding AI Agent Memory

At its core, AI Agent memory is the mechanism that enables retaining and recalling information.
This information can include conversation details, user preferences, past actions, or learned patterns.
Without memory, AI applications are often stateless, starting every interaction from scratch. This leads to a repetitive and frustrating user experience where the agent "forgets" previous context or preferences.

Why is Memory Important?

An agent's intelligence is deeply tied to its ability to recall and utilize past information. Memory enables agents to become:
CharacteristicDescription
ReflectiveLearns from past actions and outcomes
InteractiveMaintains context across continuous conversations
Proactive and ReactiveAnticipates needs or responds appropriately based on historical data
AutonomousOperates more independently using stored knowledge
The goal of implementing memory is to make agents more reliable and capable.

Types of Memory

TypeLifespanDescriptionTravel Example
Working MemorySingle taskImmediate information needed for the next step. Captures requirements, decisions, and actions"I want to book a trip to Paris" — Holds current request in immediate context
Short-term MemorySingle sessionCurrent conversation context. Can reference previous turns"Flights to Paris?" → "What about lodging there?" — Remembers that "there" is Paris
Long-term MemoryPersists across sessionsUser preferences, history, and general knowledge. The core of personalization"Ben likes skiing, prefers coffee with a mountain view, avoids advanced slopes due to injury"
Persona MemoryPersistentMaintains a consistent role and persona for the agentReinforces "expert ski planner" role — Tailors responses with expert tone and domain knowledge
Episodic MemoryPersistentStep-by-step success and failure sequences of tasks. Experience-based learningRecords a booking failure on a specific flight (no seats) → Automatically attempts alternative flights
Entity MemoryPersistentExtracts and structures people, places, and things from conversationsExtracts "Paris", "Eiffel Tower", "Le Chat Noir" → Suggests rebooking in the future
Structured RAGPersistentExtracts structured information from diverse sources. Supports precise queriesParses flight details from email → Enables answering "When was the Tuesday flight to Paris booked?"
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn', 'flowchart': {'subGraphTitleMargin': {'top': 12, 'bottom': 6}}, 'themeVariables': {'clusterBkg': '#e8efff28', 'clusterBorder': '#aabbcc'}}}%% flowchart LR Root["Memory Types"] Root --> WM["Working\nMemory"] Root --> ST["Short-term\nMemory"] Root --> LT["Long-term\nMemory"] Root --> PM["Persona\nMemory"] Root --> EM["Episodic\nMemory"] Root --> ENT["Entity\nMemory"] Root --> SR["Structured\nRAG"]

Implementing and Storing Memory

Implementing AI Agent memory involves a systematic process of memory management — including creation, storage, retrieval, consolidation, updating, and even "forgetting" (deletion).

Specialized Memory Tools

Mem0
One way to store and manage agent memory is to use specialized tools like Mem0. Mem0 acts as a persistent memory layer, allowing agents to recall relevant interactions, store user preferences and factual context, and learn from successes and failures over time.
Mem0 operates with a two-stage memory pipeline: extraction and update. First, messages added to the agent thread are sent to the Mem0 service, where an LLM summarizes conversation history to extract new memories. Afterwards, an LLM-based update step decides whether to add, modify, or delete these memories, storing them in a hybrid datastore encompassing vector, graph, and key-value databases.
Mem0: Two-Stage Memory Pipeline
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR Msg["Agent Messages"] --> Mem0["Mem0 Service"] Mem0 --> LLM["LLM\nSummarize & Extract Conversation"] LLM --> Dec{"Add / Modify\n/ Delete?"} Dec --> Store["Hybrid Datastore\nVector + Graph + Key-Value"] Store -->|"recall"| Agent["Agent"]
Cognee
Another powerful approach is Cognee — an open-source semantic memory that transforms structured and unstructured data into an embedding-based, queryable Knowledge Graph.
Cognee provides a dual-store architecture combining vector similarity search and graph relationships, enabling the agent to understand not only which information is similar, but also how concepts relate to one another.
FeatureDescription
Hybrid RetrievalBlends vector similarity, graph structure, and LLM reasoning — from chunk lookup to graph-aware QA
Living MemoryEvolves and grows while remaining queryable as a single interconnected graph
Dual-storeSupports both short-term session context and long-term persistent memory
Cognee: Dual-Store Architecture
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR Data["Structured /\nUnstructured Data"] Data -->|"cognee.add()"| VEC["Vector Store\nSimilarity Search"] Data -->|"cognee.cognify()"| GR["Graph Store\nRelations & Structure"] VEC & GR --> HQ["Hybrid Query\n(GRAPH_COMPLETION)"] HQ -->|"relevant context"| Agent["Agent"]

Storing Memory with RAG

Beyond specialized memory tools, Azure AI Search can be utilized as a memory storage and retrieval backend — particularly suited for Structured RAG.
Azure AI Search supports Structured RAG to extract and retrieve dense, structured information from large-scale datasets such as conversation history, emails, and images. It delivers "superhuman precision and recall" compared to traditional text chunking and embedding approaches.
Storing Memory with RAG (Azure AI Search)
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR D1["Conversation History"] & D2["Emails"] & D3["Images"] --> AIS["Azure AI Search"] AIS -->|"Structure & Index"| IDX["Structured Index\nDense Information Extraction"] IDX -->|"Precise Query"| Agent["Agent\nSuperhuman Precision & Recall"]

Working Memory with Sessions

setup.py
python
Working memory via sessions — by passing the same session to agent.run() calls, the agent can see the full conversation history:
working_memory.py
python
Creating a new session causes the agent to forget the previous conversation:
new_session.py
python
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR U1["User: 'Love beaches, budget $3000'"] -->|"session"| A["TravelMemoryAgent"] U2["User: 'What did I say my budget was?'"] -->|"same session"| A A -->|"Reference previous turn"| R["Accurately recalls $3000"] U3["User: 'What is my budget?'"] -->|"new_session"| B["TravelMemoryAgent\n(New session)"] B -->|"No previous conversation"| NR["Unknown — New session"]

Long-Term Memory Pattern

To remember user preferences across sessions, a persistent store outside the conversation thread is required. The agent accesses this store via tools:
memory_tools.py
python
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR subgraph Agent["MAF Agent (LLM)"] direction TB LLM["LLM"] end subgraph Tools["@tool functions"] direction TB SP["save_preference()"] GP["get_preferences()"] SH["search_hotels()"] end subgraph Store["Preference Store (dict)"] direction TB PS["user_id → [prefs]"] end LLM -->|"call"| Tools SP & GP --> PS PS -->|"persists across\nsessions"| PS SH -->|"hotel DB"| LLM

Scenario 1 — First-time user books an anniversary trip

Sarah visits for the first time. The agent stores preferences and recommends hotels using tools:
scenario1_session1.py
python
scenario1_session1_followup.py
python

Scenario 2 — Sarah returns weeks later

Sarah returns in a new thread (new session). Working memory is empty, but information remains in the long-term preference store:
scenario2_session2.py
python
scenario2_followup.py
python
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD subgraph S1["Session 1 — First Visit"] U1["Sarah: 'Romantic, spa, accessible\nBudget $700-800/night'"] U2["Sarah: 'Vegetarian, nut allergy'"] U1 & U2 -->|"save_preference()"| STORE["preference_store\n['romantic', 'spa', 'accessible',\n 'vegetarian', 'nut allergy']"] end subgraph S2["Session 2 — Returns Weeks Later"] U3["Sarah: 'Recommend a good hotel'\n(New session, no working memory)"] U3 -->|"get_preferences()"| STORE STORE -->|"Return saved preferences"| A2["TravelBookingAssistant\nPersonalized recommendations"] end

Storing Memory with Cognee Knowledge Graphs

Cognee transforms unstructured text into a queryable knowledge graph, providing agents with relation-aware long-term memory.

Part 1 — Building the Knowledge Base

We collect three types of data to build a comprehensive knowledge base:
cognee_setup.py
python
cognee_data.py
python
cognee_ingest.py
python
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD D1["Developer Profile\n(developer_intro)"] D2["Past Conversations\n(human_agent_conversations)"] D3["Python Principles\n(python_zen_principles)"] D1 & D2 -->|"cognee.add(node_set=['developer_data'])"| ADD["cognee.add()"] D3 -->|"cognee.add(node_set=['principles_data'])"| ADD ADD -->|"cognee.cognify()"| KG["Knowledge Graph\n+ Vector Embeddings"] KG -->|"cognee.memify()"| Rules["Enriched Memory\nPatterns, Rules & Relations"] KG -->|"visualize_graph()"| HTML["cognee_graph.html\nInteractive Visualization"]

Part 2 — MAF Agent with Cognee Tools

The MAF agent queries the Cognee knowledge graph via @tool functions:
cognee_tools.py
python
cognee_agent.py
python
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR subgraph MAF["MAF Agent (AgentSession)"] direction TB LLM2["CodingAssistant\n(LLM)"] Sess["create_session()\nWorking Memory"] LLM2 <--> Sess end subgraph CogneeTools["@tool functions"] direction TB SK["search_knowledge()\nGRAPH_COMPLETION"] SP2["search_principles()\nprinciples_data only"] end subgraph KG2["Cognee Knowledge Graph"] direction TB VEC["Vector Embeddings"] GR["Graph Relations"] end LLM2 -->|"call"| CogneeTools CogneeTools -->|"cognee.search()"| KG2 KG2 -->|"Return relevant knowledge"| LLM2

Making AI Agents Self-Improve

A common pattern for self-improving agents is introducing a "knowledge agent". This separate agent observes conversations between the user and the primary agent, performing the following roles:
  1. Identify valuable information: Determines which parts of the conversation are worth storing as general knowledge or specific user preferences
  2. Extract and summarize: Extracts and summarizes essential learnings or preferences from the conversation
  3. Store in a knowledge base: Stores the extracted information in a vector database or similar storage
  4. Augment future queries: When the user initiates a new query, retrieves relevant stored information and appends it to the user prompt (similar to RAG)
Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD User["User"] -->|"New query"| KA["Knowledge Agent\n(Observer)"] KA -->|"Retrieve relevant knowledge\n(RAG)"| VDB["Vector DB\nKnowledge Base"] VDB -->|"Add stored context"| PA["Primary Agent"] PA -->|"Response"| User PA -->|"Observe conversation"| KA KA -->|"Valuable info\nExtract, summarize & store"| VDB

Optimizations for Memory

OptimizationDescription
Latency ManagementQuickly verify the value of storing or retrieving information using cheaper, faster models first, calling complex extraction and retrieval processes only when needed
Knowledge Base MaintenanceManage costs by moving less frequently used information to "cold storage" as the knowledge base grows

Summary

Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR Root["Memory for\nAI Agents"] Root --> Types["Memory Types"] Root --> Tools["Memory Tools"] Root --> Pattern["Self-Improving\nPattern"] Types --> T1["Working\ncreate_session()"] Types --> T2["Short-term\nIn-session context"] Types --> T3["Long-term\n@tool persistent store"] Types --> T4["Persona·Episodic\nEntity·Structured RAG"] Tools --> TL1["Mem0\nExtract + Update Pipeline"] Tools --> TL2["Cognee\nKnowledge Graph + Vector"] Tools --> TL3["Azure AI Search\nStructured RAG Backend"] Pattern --> P1["Knowledge Agent\nObserve & Extract Conversation"] Pattern --> P2["RAG Augmentation\nAdd context to future queries"]
Memory TypeMAF MechanismLifespan
Workingagent.create_session()Single conversation
Short-termCumulative context within threadSingle task/session
Long-termExternal store accessed via @tool functionsPersists across sessions
  • Provides working memory via agent.create_session() — the agent sees the full conversation history within the session
  • New sessions lose context — without long-term memory, the agent cannot recall past conversations
  • @tool functions bridge the gap — allowing the agent to store and retrieve information from a persistent store
  • Mem0: Transitions stateless agents into stateful ones using a two-stage pipeline (extraction → update)
  • Cognee: Relation-aware long-term memory via knowledge graphs + vector embeddings — understands relationships between concepts using GRAPH_COMPLETION search
  • Forms a self-improving loop where the more preferences and history are stored, the more refined the agent's recommendations become
Jooojub
System S/W engineer
Explore Tags
Series
    Recent Post
    © 2026. jooojub. All right reserved.