AI Agents for Beginners - 12. Context Engineering
How Context Engineering differs from Prompt Engineering, the five types of context in AI Agents, practical management strategies, and how to recognize and fix Context Poisoning, Distraction, Confusion, and Clash failures.
June 9, 2026
AI Agents for Beginners - 12. Context Engineering
This article summarizes Lesson 12 of Microsoft's AI Agents for Beginners course.
What is Context Engineering?
In AI Agents, Context refers to all the information the agent needs when planning its next action.
Context Engineering is the practice of ensuring an agent is equipped with the right information needed to complete each step.
Because the context window is limited in size, agent developers must design systems and processes to add, remove, and compress information within the context window.
Context Engineering is the practice of ensuring an agent is equipped with the right information needed to complete each step.
Because the context window is limited in size, agent developers must design systems and processes to add, remove, and compress information within the context window.
Prompt Engineering vs Context Engineering
| Category | Prompt Engineering | Context Engineering |
|---|---|---|
| Focus | Single static instruction | Managing a dynamic collection of information |
| Goal | Effectively guide the AI agent with rules | Ensure the agent has what it needs over time |
| Nature | One-time optimization | A repeatable and reliable process |
The core of Context Engineering lies in making this process repeatable and reliable.
Types of Context
The context an AI agent must manage comes from a variety of sources rather than a single origin:
| Type | Description |
|---|---|
| Instructions | The agent's "rules" — prompts, system messages, few-shot examples, and descriptions of available tools |
| Knowledge | Facts and information retrieved from databases, and long-term memory accumulated by the agent. Includes RAG systems |
| Tools | Definitions of external functions, APIs, and MCP servers, as well as feedback (results) after execution |
| Conversation History | Ongoing conversation with the user. Grows over time and consumes context window space |
| User Preferences | User preferences learned over time. Retrieved and referenced when making key decisions |
Context Types in AI AgentsMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD Agent["AI Agent\nContext Window"] Agent --> I["Instructions\nPrompts, System Messages, Few-shot, Tool Descriptions"] Agent --> K["Knowledge\nRAG, Databases, Long-term Memory"] Agent --> T["Tools\nAPIs, MCP Servers, Function Results"] Agent --> CH["Conversation History\nMulti-turn Dialogue History"] Agent --> UP["User Preferences\nPreferences, Habits, Feedback"]
Strategies for Effective Context Engineering
Good context engineering begins with solid planning:
- Define Clear Results — Clearly define the desired outcomes of the task assigned to the AI Agent. You must answer the question: "What does the world look like when the agent is done?"
- Map the Context — After defining results, determine "What information does the agent need to complete this task?" and map out where that information resides.
- Create Context Pipelines — Once you know where the information is located, decide "How will the agent obtain this information?" Leverage RAG, MCP servers, and other tools.
Context Engineering WorkflowMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR R["Define Clear Results\nClearly define outcomes"] --> M["Map the Context\nMap required info locations"] M --> P["Create Context Pipelines\nDeliver info via RAG, MCP, Tools"] P --> Agent["AI Agent\nEquipped with proper context"]
Planning Strategies
Managing Context
Once information begins flowing into the context window, it must be actively managed:
| Strategy | Description | When to Use |
|---|---|---|
| Agent Scratchpad | Stores relevant information for the current session in an external file/object outside the context window | Preserving intermediate results within a single session |
| Memories | Stores and retrieves relevant information across multiple sessions | Long-term retention of user preferences and summaries |
| Compressing Context | Removes older messages through summarization and trimming | Just before reaching context window limits |
| Multi-Agent Systems | Each agent maintains an independent context window | Distributing complex parallel tasks |
| Sandbox Environments | Runs code execution and heavy processing in an isolated environment, reading back only results | Large document processing and code execution |
| Runtime State Objects | Containers that store step-by-step results of subtasks for complex tasks | Multi-step complex workflows |
Context Management StrategiesMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD CW["Context Window\n(Limited Size)"] CW -->|"Temporary single-session storage"| SP["Agent Scratchpad\nFiles / Runtime objects"] CW -->|"Long-term cross-session retention"| MEM["Memories\nSummaries, Preferences, Feedback"] CW -->|"When approaching limits"| COMP["Compressing Context\nSummarization / Trimming"] CW -->|"Complex parallel tasks"| MAS["Multi-Agent Systems\nDistributed independent contexts"] CW -->|"Code / Large-scale processing"| SBX["Sandbox Environments\nReturn results only"] CW -->|"Multi-step workflows"| RSO["Runtime State Objects\nSubtask result containers"]
Example of Context Engineering
Comparing the difference between the two approaches using the request: "Book me a trip to Paris."
| Approach | Behavior | Example Response |
|---|---|---|
| Prompt Engineering only | Handles only the immediate prompt | "Okay, when would you like to go to Paris?" |
| Context Engineering | Checks the calendar, looks up preferences from long-term memory, identifies booking tools, and responds | "Hey! I see you're free the first week of October. Shall I look for direct flights to Paris on [Preferred Airline] within your usual budget of [Budget]?" |
An agent with Context Engineering applied automatically performs the following before responding:
- Checks available dates in the calendar in real time
- Recalls preferred airline, budget, and direct flight preference from long-term memory
- Identifies the tools needed to book flights and hotels
Common Context Failures
Context Poisoning
Definition: A phenomenon where an LLM-generated hallucination or error enters the context and is repeatedly referenced, causing the agent to pursue impossible goals or develop meaningless strategies.
Travel Example: The agent hallucinates a direct flight from a small regional airport to an international city. This nonexistent flight is stored in the context, and when later asked to book, the agent repeatedly attempts to find tickets for this impossible route.
Solution: Implement a step to verify flight existence and routes using a real-time API before adding flight details to the context. If validation fails, "quarantine" the faulty information so it is not used going forward.
Context Poisoning MitigationMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD LLM["LLM Response\n(Includes flight info)"] --> V{"Real-time API Verification"} V -->|"Verification successful"| CTX["Add to context\nProceed normally"] V -->|"Verification failed\nHallucination detected"| Q["Quarantine\nBlock context"] Q --> Fresh["Fresh context thread\nPrevent contamination"]
Context Distraction
Definition: A phenomenon where the context becomes excessively large, leading the model to over-focus on the accumulated history rather than what it learned during training, resulting in repetitive or unproductive behavior. Errors can occur even before the context window is full.
Travel Example: While discussing various destinations over an extended period, you thoroughly explained a backpacking trip from two years ago. Later, when you ask "Find cheap flights for next month," the agent gets bogged down in backpacking gear or past itineraries and fails to address the current request.
Solution: Once a conversation reaches a certain number of turns or the context grows too large, the agent summarizes the most recent relevant parts and uses this compressed summary for subsequent LLM calls.
Context Distraction SummarizationMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR Long["Long Conversation History\n(Includes irrelevant content)"] -->|"Threshold exceeded"| Sum["Context Summarization\nCurrent goals / relevant details only"] Sum --> Next["Next LLM Call\nUse compressed summary"] Next --> Focused["Focused Response\nTailored to current request"]
Context Confusion
Definition: A phenomenon where unnecessary context—especially too many tools—causes the model to produce erroneous responses or invoke irrelevant tools. Smaller models are especially prone to this.
Travel Example: The agent has access to dozens of tools such as
book_flight, book_hotel, rent_car, find_tours, currency_converter, weather_forecast, and restaurant_reservations. When asked "What's the best way to get around Paris?", overlapping or difficult-to-distinguish tool descriptions lead the agent to call book_flight within Paris or attempt rent_car.Solution: RAG over tool descriptions — Dynamically select and provide only the most relevant tools to the LLM based on the user's query. It is recommended to limit tool selection to fewer than 30 tools.
Context Confusion RAG FilteringMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD Query["User Query\n'How to get around Paris?'"] --> RAG["RAG over Tool Descriptions\nSelect relevant tools from vector DB"] RAG --> Focused["Focused Tool Set\nrent_car, public_transport_info"] Focused --> LLM["LLM\nInvoke appropriate tools"] LLM --> Response["Accurate Response"]
Context Clash
Definition: A phenomenon where conflicting information exists within the context, leading to inconsistent reasoning or incorrect final responses. This frequently occurs when information arrives incrementally and early incorrect assumptions linger in the context.
Travel Example: Initially, you tell the agent, "I want to fly economy class." Later, you change your mind: "Let's go business class for this trip." If both instructions remain in the context, the agent becomes confused about which preference to prioritize.
Solution: Context Pruning — Remove or explicitly override older instructions when new instructions contradict them. Alternatively, use a Scratchpad to reconcile conflicting preferences before making a decision.
Context Clash ResolutionMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD Old["Previous Instruction\n'Economy class'"] --> Conflict{"Conflict Detected\nNew instruction arrives"} New["New Instruction\n'Business class'"] --> Conflict Conflict -->|"Context Pruning"| Prune["Remove previous instruction\nKeep new instruction only"] Conflict -->|"Scratchpad"| Scratch["Reconcile in Scratchpad\nDetermine final consistent instruction"] Prune & Scratch --> Clean["Consistent Context\nAgent operates normally"]
Chat History Reduction with Agent Scratchpad
Every LLM has a finite context window, representing the maximum number of tokens it can process in a single request. As a multi-turn conversation progresses:
- Token counts grow linearly with each user message and assistant response
- Resending the full conversation history every time causes prompt tokens to account for the majority of costs
- Eventually, if the conversation exceeds the context window, the model truncates it or returns an error
| Strategy | How It Works | Trade-offs |
|---|---|---|
| Truncation | Deletes the oldest messages | Loss of initial context |
| Summarization | Compresses older messages into a summary | Retains key points, but loses some details |
| Scratchpad / External Memory | Stores essential facts outside the conversation | Requires tool calls, but survives any reduction |
Creating a Context-Aware Agent
context_aware_agent.pypython
Simulating a Long Conversation
Observe how context accumulates through a multi-turn conversation. The agent must preserve key details (preferences, budget, travel dates) across multiple turns:
long_conversation.pypython
Context Growth in Long ConversationsMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD T1["Turn 1\nJapan, Sushi, Temples, Photography"] --> T2["Turn 2\nBudget $3000, Solo, 10 days, April"] T2 --> T3["Turn 3\nTest context retention"] T3 --> T4["Turn 4\nAccommodation: Traditional Japanese Ryokan"] T4 --> T5["Turn 5\nDate change: October autumn foliage"] T5 --> T6["Turn 6\nRequest full summary"] T1 & T2 & T4 & T5 -->|"Accumulated context"| CTX["Context Window\n(Growing)"] CTX -->|"Short conversation: Normal"| T6
Context Summarization Pattern
When conversations grow long, use a Summarization Tool to compress accumulated context.
This pattern forms the foundation for more sophisticated history reduction:
- The agent identifies key facts from the conversation
- Calls the summarization tool to store them persistently
- Because the summary captures core essentials, older messages can be safely pruned
summarization_tool.pypython
summarization_demo.pypython
Summarization and Scratchpad PatternMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD Conv["Multi-turn Conversation\nAccumulated preferences, budget, dates, interests"] --> Agent["SummarizingTravelAgent"] Agent -->|"Identify key facts"| Tool["summarize_preferences()\nStore compact summary"] Tool --> Sum["[SUMMARY] User preferences recorded: ..."] Sum -->|"Older messages can be pruned"| Trim["Context Window Reduction"] Trim --> Next["Next LLM Call\nSummary-based response"] Next -->|"Continuous continuity"| User["User\nResponse retaining previous context"]
Agent Scratchpad Example — Notes persisted across sessions in an external file:
vacation_agent_scratchpad.mdmarkdown
Because the Scratchpad resides outside the context window, it survives any form of context reduction and provides cross-session continuity.
Summary
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR Root["Context\nEngineering"] Root --> PE["vs Prompt Eng.\nStatic vs Dynamic"] Root --> Types["Types of Context\n5 Sources"] Root --> Plan["Planning\nStrategies"] Root --> Mgmt["Managing\nContext\n6 Strategies"] Root --> Fail["Common\nFailures"] Types --> I["Instructions"] Types --> K["Knowledge"] Types --> T["Tools"] Types --> CH["Conversation\nHistory"] Types --> UP["User Preferences"] Plan --> P1["Define Clear Results"] Plan --> P2["Map the Context"] Plan --> P3["Create Pipelines"] Mgmt --> M1["Scratchpad"] Mgmt --> M2["Memories"] Mgmt --> M3["Compression"] Mgmt --> M4["Multi-Agent"] Mgmt --> M5["Sandbox"] Mgmt --> M6["Runtime State"] Fail --> F1["Poisoning\nHallucination propagation"] Fail --> F2["Distraction\nOver-focus on history"] Fail --> F3["Confusion\nTool overload"] Fail --> F4["Clash\nInformation conflict"]
- Context Engineering goes beyond Prompt Engineering to build a repeatable and reliable process ensuring agents are equipped with the dynamic information they need over time
- Context is composed of five types: Instructions, Knowledge, Tools, Conversation History, and User Preferences
- Effectively managed through Planning Strategies (Define Results → Map Context → Create Pipelines) and six practical strategies (Scratchpad, Memories, Compression, Multi-Agent, Sandbox, Runtime State)
- Recognize four Context Failures—Context Poisoning, Distraction, Confusion, and Clash—and mitigate them via verification, summarization, RAG, and pruning
- Maintain context continuity while reducing token costs even across extended conversations using the Chat Summarization + Agent Scratchpad pattern