AI Agents for Beginners - 12. Context Engineering

How Context Engineering differs from Prompt Engineering, the five types of context in AI Agents, practical management strategies, and how to recognize and fix Context Poisoning, Distraction, Confusion, and Clash failures.
June 9, 2026

AI Agents for Beginners - 12. Context Engineering

This article summarizes Lesson 12 of Microsoft's AI Agents for Beginners course.

What is Context Engineering?

In AI Agents, Context refers to all the information the agent needs when planning its next action.
Context Engineering is the practice of ensuring an agent is equipped with the right information needed to complete each step.
Because the context window is limited in size, agent developers must design systems and processes to add, remove, and compress information within the context window.

Prompt Engineering vs Context Engineering

CategoryPrompt EngineeringContext Engineering
FocusSingle static instructionManaging a dynamic collection of information
GoalEffectively guide the AI agent with rulesEnsure the agent has what it needs over time
NatureOne-time optimizationA repeatable and reliable process
The core of Context Engineering lies in making this process repeatable and reliable.

Types of Context

The context an AI agent must manage comes from a variety of sources rather than a single origin:
TypeDescription
InstructionsThe agent's "rules" — prompts, system messages, few-shot examples, and descriptions of available tools
KnowledgeFacts and information retrieved from databases, and long-term memory accumulated by the agent. Includes RAG systems
ToolsDefinitions of external functions, APIs, and MCP servers, as well as feedback (results) after execution
Conversation HistoryOngoing conversation with the user. Grows over time and consumes context window space
User PreferencesUser preferences learned over time. Retrieved and referenced when making key decisions
Context Types in AI Agents
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD Agent["AI Agent\nContext Window"] Agent --> I["Instructions\nPrompts, System Messages, Few-shot, Tool Descriptions"] Agent --> K["Knowledge\nRAG, Databases, Long-term Memory"] Agent --> T["Tools\nAPIs, MCP Servers, Function Results"] Agent --> CH["Conversation History\nMulti-turn Dialogue History"] Agent --> UP["User Preferences\nPreferences, Habits, Feedback"]

Strategies for Effective Context Engineering

Good context engineering begins with solid planning:
  1. Define Clear Results — Clearly define the desired outcomes of the task assigned to the AI Agent. You must answer the question: "What does the world look like when the agent is done?"
  2. Map the Context — After defining results, determine "What information does the agent need to complete this task?" and map out where that information resides.
  3. Create Context Pipelines — Once you know where the information is located, decide "How will the agent obtain this information?" Leverage RAG, MCP servers, and other tools.
Context Engineering Workflow
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR R["Define Clear Results\nClearly define outcomes"] --> M["Map the Context\nMap required info locations"] M --> P["Create Context Pipelines\nDeliver info via RAG, MCP, Tools"] P --> Agent["AI Agent\nEquipped with proper context"]

Planning Strategies

Managing Context

Once information begins flowing into the context window, it must be actively managed:
StrategyDescriptionWhen to Use
Agent ScratchpadStores relevant information for the current session in an external file/object outside the context windowPreserving intermediate results within a single session
MemoriesStores and retrieves relevant information across multiple sessionsLong-term retention of user preferences and summaries
Compressing ContextRemoves older messages through summarization and trimmingJust before reaching context window limits
Multi-Agent SystemsEach agent maintains an independent context windowDistributing complex parallel tasks
Sandbox EnvironmentsRuns code execution and heavy processing in an isolated environment, reading back only resultsLarge document processing and code execution
Runtime State ObjectsContainers that store step-by-step results of subtasks for complex tasksMulti-step complex workflows
Context Management Strategies
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD CW["Context Window\n(Limited Size)"] CW -->|"Temporary single-session storage"| SP["Agent Scratchpad\nFiles / Runtime objects"] CW -->|"Long-term cross-session retention"| MEM["Memories\nSummaries, Preferences, Feedback"] CW -->|"When approaching limits"| COMP["Compressing Context\nSummarization / Trimming"] CW -->|"Complex parallel tasks"| MAS["Multi-Agent Systems\nDistributed independent contexts"] CW -->|"Code / Large-scale processing"| SBX["Sandbox Environments\nReturn results only"] CW -->|"Multi-step workflows"| RSO["Runtime State Objects\nSubtask result containers"]

Example of Context Engineering

Comparing the difference between the two approaches using the request: "Book me a trip to Paris."
ApproachBehaviorExample Response
Prompt Engineering onlyHandles only the immediate prompt"Okay, when would you like to go to Paris?"
Context EngineeringChecks the calendar, looks up preferences from long-term memory, identifies booking tools, and responds"Hey! I see you're free the first week of October. Shall I look for direct flights to Paris on [Preferred Airline] within your usual budget of [Budget]?"
An agent with Context Engineering applied automatically performs the following before responding:
  • Checks available dates in the calendar in real time
  • Recalls preferred airline, budget, and direct flight preference from long-term memory
  • Identifies the tools needed to book flights and hotels

Common Context Failures

Context Poisoning

Definition: A phenomenon where an LLM-generated hallucination or error enters the context and is repeatedly referenced, causing the agent to pursue impossible goals or develop meaningless strategies.
Travel Example: The agent hallucinates a direct flight from a small regional airport to an international city. This nonexistent flight is stored in the context, and when later asked to book, the agent repeatedly attempts to find tickets for this impossible route.
Solution: Implement a step to verify flight existence and routes using a real-time API before adding flight details to the context. If validation fails, "quarantine" the faulty information so it is not used going forward.
Context Poisoning Mitigation
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD LLM["LLM Response\n(Includes flight info)"] --> V{"Real-time API Verification"} V -->|"Verification successful"| CTX["Add to context\nProceed normally"] V -->|"Verification failed\nHallucination detected"| Q["Quarantine\nBlock context"] Q --> Fresh["Fresh context thread\nPrevent contamination"]

Context Distraction

Definition: A phenomenon where the context becomes excessively large, leading the model to over-focus on the accumulated history rather than what it learned during training, resulting in repetitive or unproductive behavior. Errors can occur even before the context window is full.
Travel Example: While discussing various destinations over an extended period, you thoroughly explained a backpacking trip from two years ago. Later, when you ask "Find cheap flights for next month," the agent gets bogged down in backpacking gear or past itineraries and fails to address the current request.
Solution: Once a conversation reaches a certain number of turns or the context grows too large, the agent summarizes the most recent relevant parts and uses this compressed summary for subsequent LLM calls.
Context Distraction Summarization
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR Long["Long Conversation History\n(Includes irrelevant content)"] -->|"Threshold exceeded"| Sum["Context Summarization\nCurrent goals / relevant details only"] Sum --> Next["Next LLM Call\nUse compressed summary"] Next --> Focused["Focused Response\nTailored to current request"]

Context Confusion

Definition: A phenomenon where unnecessary context—especially too many tools—causes the model to produce erroneous responses or invoke irrelevant tools. Smaller models are especially prone to this.
Travel Example: The agent has access to dozens of tools such as book_flight, book_hotel, rent_car, find_tours, currency_converter, weather_forecast, and restaurant_reservations. When asked "What's the best way to get around Paris?", overlapping or difficult-to-distinguish tool descriptions lead the agent to call book_flight within Paris or attempt rent_car.
Solution: RAG over tool descriptions — Dynamically select and provide only the most relevant tools to the LLM based on the user's query. It is recommended to limit tool selection to fewer than 30 tools.
Context Confusion RAG Filtering
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD Query["User Query\n'How to get around Paris?'"] --> RAG["RAG over Tool Descriptions\nSelect relevant tools from vector DB"] RAG --> Focused["Focused Tool Set\nrent_car, public_transport_info"] Focused --> LLM["LLM\nInvoke appropriate tools"] LLM --> Response["Accurate Response"]

Context Clash

Definition: A phenomenon where conflicting information exists within the context, leading to inconsistent reasoning or incorrect final responses. This frequently occurs when information arrives incrementally and early incorrect assumptions linger in the context.
Travel Example: Initially, you tell the agent, "I want to fly economy class." Later, you change your mind: "Let's go business class for this trip." If both instructions remain in the context, the agent becomes confused about which preference to prioritize.
Solution: Context Pruning — Remove or explicitly override older instructions when new instructions contradict them. Alternatively, use a Scratchpad to reconcile conflicting preferences before making a decision.
Context Clash Resolution
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD Old["Previous Instruction\n'Economy class'"] --> Conflict{"Conflict Detected\nNew instruction arrives"} New["New Instruction\n'Business class'"] --> Conflict Conflict -->|"Context Pruning"| Prune["Remove previous instruction\nKeep new instruction only"] Conflict -->|"Scratchpad"| Scratch["Reconcile in Scratchpad\nDetermine final consistent instruction"] Prune & Scratch --> Clean["Consistent Context\nAgent operates normally"]

Chat History Reduction with Agent Scratchpad

Every LLM has a finite context window, representing the maximum number of tokens it can process in a single request. As a multi-turn conversation progresses:
  • Token counts grow linearly with each user message and assistant response
  • Resending the full conversation history every time causes prompt tokens to account for the majority of costs
  • Eventually, if the conversation exceeds the context window, the model truncates it or returns an error
StrategyHow It WorksTrade-offs
TruncationDeletes the oldest messagesLoss of initial context
SummarizationCompresses older messages into a summaryRetains key points, but loses some details
Scratchpad / External MemoryStores essential facts outside the conversationRequires tool calls, but survives any reduction

Creating a Context-Aware Agent

context_aware_agent.py
python

Simulating a Long Conversation

Observe how context accumulates through a multi-turn conversation. The agent must preserve key details (preferences, budget, travel dates) across multiple turns:
long_conversation.py
python
Context Growth in Long Conversations
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD T1["Turn 1\nJapan, Sushi, Temples, Photography"] --> T2["Turn 2\nBudget $3000, Solo, 10 days, April"] T2 --> T3["Turn 3\nTest context retention"] T3 --> T4["Turn 4\nAccommodation: Traditional Japanese Ryokan"] T4 --> T5["Turn 5\nDate change: October autumn foliage"] T5 --> T6["Turn 6\nRequest full summary"] T1 & T2 & T4 & T5 -->|"Accumulated context"| CTX["Context Window\n(Growing)"] CTX -->|"Short conversation: Normal"| T6

Context Summarization Pattern

When conversations grow long, use a Summarization Tool to compress accumulated context.
This pattern forms the foundation for more sophisticated history reduction:
  1. The agent identifies key facts from the conversation
  2. Calls the summarization tool to store them persistently
  3. Because the summary captures core essentials, older messages can be safely pruned
summarization_tool.py
python
summarization_demo.py
python
Summarization and Scratchpad Pattern
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD Conv["Multi-turn Conversation\nAccumulated preferences, budget, dates, interests"] --> Agent["SummarizingTravelAgent"] Agent -->|"Identify key facts"| Tool["summarize_preferences()\nStore compact summary"] Tool --> Sum["[SUMMARY] User preferences recorded: ..."] Sum -->|"Older messages can be pruned"| Trim["Context Window Reduction"] Trim --> Next["Next LLM Call\nSummary-based response"] Next -->|"Continuous continuity"| User["User\nResponse retaining previous context"]
Agent Scratchpad Example — Notes persisted across sessions in an external file:
vacation_agent_scratchpad.md
markdown
Because the Scratchpad resides outside the context window, it survives any form of context reduction and provides cross-session continuity.

Summary

Process Flowchart
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart LR Root["Context\nEngineering"] Root --> PE["vs Prompt Eng.\nStatic vs Dynamic"] Root --> Types["Types of Context\n5 Sources"] Root --> Plan["Planning\nStrategies"] Root --> Mgmt["Managing\nContext\n6 Strategies"] Root --> Fail["Common\nFailures"] Types --> I["Instructions"] Types --> K["Knowledge"] Types --> T["Tools"] Types --> CH["Conversation\nHistory"] Types --> UP["User Preferences"] Plan --> P1["Define Clear Results"] Plan --> P2["Map the Context"] Plan --> P3["Create Pipelines"] Mgmt --> M1["Scratchpad"] Mgmt --> M2["Memories"] Mgmt --> M3["Compression"] Mgmt --> M4["Multi-Agent"] Mgmt --> M5["Sandbox"] Mgmt --> M6["Runtime State"] Fail --> F1["Poisoning\nHallucination propagation"] Fail --> F2["Distraction\nOver-focus on history"] Fail --> F3["Confusion\nTool overload"] Fail --> F4["Clash\nInformation conflict"]
  • Context Engineering goes beyond Prompt Engineering to build a repeatable and reliable process ensuring agents are equipped with the dynamic information they need over time
  • Context is composed of five types: Instructions, Knowledge, Tools, Conversation History, and User Preferences
  • Effectively managed through Planning Strategies (Define Results → Map Context → Create Pipelines) and six practical strategies (Scratchpad, Memories, Compression, Multi-Agent, Sandbox, Runtime State)
  • Recognize four Context Failures—Context Poisoning, Distraction, Confusion, and Clash—and mitigate them via verification, summarization, RAG, and pruning
  • Maintain context continuity while reducing token costs even across extended conversations using the Chat Summarization + Agent Scratchpad pattern

Jooojub
System S/W engineer
Explore Tags
Series
    Recent Post
    © 2026. jooojub. All right reserved.