AI Agents for Beginners - 9. Metacognition in AI Agents
Metacognition in AI agents — self-reflection, corrective RAG, pre-emptive context loading, goal bootstrapping, intent-aware search, code generation, and SQL as RAG.
June 9, 2026
AI Agents for Beginners - 9. Metacognition in AI Agents
This post summarizes Lesson 09 from Microsoft's AI Agents for Beginners course.
What is Metacognition?
Metacognition is "thinking about thinking."
In AI agents, metacognition refers to the ability of an agent to recognize its own internal processes and monitor, regulate, and adapt its behavior.
In AI agents, metacognition refers to the ability of an agent to recognize its own internal processes and monitor, regulate, and adapt its behavior.
Key challenges addressed by metacognition:
| Challenge | Description |
|---|---|
| Transparency | The agent must be able to explain its reasoning and decisions |
| Reasoning | Enhances the ability to synthesize information and make sound decisions |
| Adaptation | Adapts to new environments and changing conditions |
| Perception | Improves accuracy in perceiving and interpreting environmental data |
Why metacognition is important in agent design:
Importance of MetacognitionMermaidflowchart LR MC["Metacognition\nRecognizing Own Thought Process"] MC --> SR["Self-Reflection\nEvaluating Performance & Identifying Improvement Areas"] MC --> AD["Adaptability\nAdjusting Strategy Based on Past Experience"] MC --> EC["Error Correction\nAutonomous Error Detection & Correction"] MC --> RM["Resource Management\nOptimizing Time & Compute Resources"]
Core examples of metacognition:
- "I prioritized cheap flights... but I might have missed direct flights. Let me check again."
- "Whenever the user said 'it's too complicated,' the issue was that I recommended items based on popularity."
The agent components—Persona, Tools, and Skills—combine to form an "expertise unit" that executes specific tasks, and metacognition enables this unit to improve on its own.
Planning in Agents
Planning is a core element of AI agent behavior that designs steps to achieve a goal while taking into account the current state, resources, and obstacles.
9-step planning process of a travel agent:
Travel Agent Planning FlowMermaidflowchart TD P1["1. Gather Preferences\nCollect dates, budget, interests"] --> P2["2. Retrieve Information\nSearch flights, hotels, attractions"] --> P3["3. Generate Recommendations\nCreate personalized itinerary"] --> P4["4. Present to User\nShare itinerary"] --> P5["5. Collect Feedback\nGather user feedback"] --> P6["6. Adjust Based on Feedback\nUpdate itinerary"] --> P7["7. Final Confirmation\nFinal review"] --> P8["8. Book & Confirm\nProceed with booking"] --> P9["9. Ongoing Support\nSupport during trip"]
Travel Agent implementation applying metacognition:
travel_agent_metacognition.pypython
Metacognition LoopMermaidflowchart TD G["gather_preferences()\nInitialize preferences"] --> R["generate_recommendations()\nSearch flights/hotels/attractions + Generate itinerary"] --> U["Present itinerary to user"] --> FB["User Feedback\n{liked, disliked}"] FB --> Adj["adjust_based_on_feedback()\nAccumulate experience_data\nReadjust preferences"] Adj -->|"Updated Preferences"| R
Accumulating feedback in
experience_data, readjusting preferences with adjust_based_on_feedback, and generating a new itinerary forms the core of the metacognition loop.Corrective RAG System
Comparing two information access strategies:
| Corrective RAG | Pre-emptive Context Load | |
|---|---|---|
| Approach | Retrieve information from external sources at query time | Pre-load relevant context prior to processing |
| Updates | Real-time information | Pre-loaded information |
| Speed | Incurs retrieval latency | Immediate response |
| Accuracy | Feedback-based error correction | Fixed context |
| Best For | Dynamic data, when error correction is needed | Known domain, when fast responses are required |
RAG Tool Calling vs Pre-Emptive Context LoadMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD subgraph RAG["RAG Tool Calling"] direction TD LLM1["LLM"] LLM1 --> T1["Tool"] LLM1 --> T2["Tool"] LLM1 --> T3["Tool"] T1 & T2 & T3 --> R1["Response"] end subgraph PCL["Pre-Emptive Context Load"] direction TD LLM2["LLM"] LLM2 --> CTX["Context\n(Provided before query processing)"] CTX --> R2["Response"] end
Corrective RAG Approach
Corrective RAG corrects errors and improves accuracy through three components:
| Element | Description |
|---|---|
| Prompting Technique | Crafting specific prompts that guide the agent to retrieve relevant information |
| Tool | Implementing algorithms and mechanisms to evaluate the relevance of retrieved information and generate accurate responses |
| Evaluation | Continuously evaluating agent performance and iterating on adjustments to improve accuracy and efficiency |
Step-by-Step Implementation of Corrective RAG
Five steps to applying Corrective RAG to a Travel Agent:
Step 1 — Gather User Input
Collect travel preferences such as destination, dates, budget, and interests.
step1_preferences.pypython
Step 2 — Retrieve Information and Generate Initial Recommendations
Search for flights, hotels, and attractions based on the collected preferences and generate an initial itinerary.
step2_retrieve.pypython
Step 3 — Collect User Feedback
Collect positive and negative user feedback on the initial recommendations.
step3_feedback.pypython
Step 4 — Apply Corrective RAG
- Prompting Technique: Formulate a new search query reflecting the feedback
step4_prompting.pypython
- Tool: Re-search with updated preferences and generate a new itinerary
step4_tool.pypython
- Evaluation: Continuous evaluation function reflecting feedback into preferences
step4_evaluation.pypython
Step 5 — Putting It Together: Corrective RAG Travel Agent
This final example integrates the above steps into a single class.
adjust_based_on_feedback handles feedback accumulation → preference update → new itinerary generation all in one go.corrective_rag_travel_agent.pypython
Corrective RAG Travel Agent FlowMermaidflowchart TD GP["gather_preferences()\nStore destination, budget, interests"] --> GR["generate_recommendations()\nCall retrieve_information()"] GR --> RT["search_flights + hotels + attractions"] RT --> CI["create_itinerary()"] CI --> SU["Present itinerary to user"] SU --> FB["User Feedback\n{liked, disliked}"] FB --> AF["adjust_based_on_feedback()\nexperience_data.append(feedback)\nadjust_preferences() → Update preferences"] AF -->|"Updated user_preferences"| GR
Corrective RAG Flow
Corrective RAG FlowMermaidflowchart TD User["User Input\nDestination, budget, interests"] --> Retrieve["Retrieve Information\nFlights, hotels, attractions"] --> Gen["Generate Initial Itinerary"] --> Feedback["Collect User Feedback"] Feedback --> Adjust{Analyze Feedback} Adjust -->|"disliked items"| NewQuery["Rewrite Query\nAdd preferences['avoid']"] NewQuery --> Retrieve Adjust -->|"Satisfied"| Done["Finalize Itinerary"]
Pre-emptive Context Load
Pre-loads known domain knowledge into memory before processing to respond immediately without external retrieval:
preemptive_context.pypython
Pre-emptive Context Load FlowMermaidflowchart TD Init["__init__()\nPre-load self.context\n{Paris, Tokyo, NY, Sydney}"] --> Query["User destination query"] --> Get["get_destination_info(destination)\nself.context.get(destination)"] Get --> Check{Context exists?} Check -->|"Yes"| Hit["country, currency, language, attractions\nReturn immediately (no external retrieval)"] Check -->|"No"| Miss["Return 'No information available'"]
Bootstrapping a Plan with a Goal Before Iterating
An approach where a goal is defined upfront, followed by iterative refinement of the plan.
Initial planning is established with bootstrap_plan, and progressively optimized using iterate_plan:
Initial planning is established with bootstrap_plan, and progressively optimized using iterate_plan:
bootstrap_plan.pypython
Goal Bootstrapping LoopMermaidflowchart TD Goal["Define Clear Goal\n(e.g., maximize satisfaction, budget ≤ $2000)"] --> Bootstrap["bootstrap_plan\nCreate initial plan based on goal & budget"] --> Execute["Execute Plan"] --> Eval{Evaluate Results} Eval -->|"Can be improved"| Iterate["iterate_plan\nReplace with better destinations"] Iterate --> Execute Eval -->|"Optimal"| Done["Finalize Travel Plan"]
Taking advantage of LLM for Re-ranking and Scoring
LLM Re-ranking is a technique where instead of using initial retrieval results directly, an LLM evaluates the relevance and quality of each candidate and reorders them into an optimal sequence.
| Stage | Description |
|---|---|
| Retrieval | Initial retrieval of candidate documents/responses based on the query |
| Re-ranking | LLM reorders candidates by relevance and quality — placing the most helpful information at the top |
| Scoring | LLM assigns scores to each candidate to select the optimal response or document |
Example of an LLM re-ranking and scoring destinations based on user preferences in a travel agent:
llm_reranking.pypython
LLM Re-ranking & Scoring FlowMermaidflowchart TD Query["User Preferences\nactivity: sightseeing\nculture: diverse"] --> Retrieve["Initial Candidate List\n(Paris, Tokyo, NYC, Sydney)"] --> Prompt["generate_prompt()\nInclude preferences + destination descriptions"] --> LLM["LLM (Azure OpenAI)\nEvaluate relevance & quality"] LLM --> Rerank["Re-ranking\nSort best-matching destinations to the top"] Rerank --> Score["Scoring\nAssign scores to each candidate"] Score --> Result["Final Recommendations\n(Including ranks & scores)"]
RAG: Prompting Technique vs Tool
RAG can be utilized in two ways: as a prompting technique or as an integrated tool:
| Aspect | Prompting Technique | Tool |
|---|---|---|
| Approach | Manually crafting prompts for each query | Automating retrieval and generation |
| Control | Fine-grained control over the retrieval process | Full end-to-end automated handling |
| Flexibility | Custom prompts tailored to specific requirements | More efficient for large-scale implementations |
| Complexity | Requires writing and tuning prompts | Easy to integrate into AI Agent architectures |
rag_prompting_vs_tool.pypython
Evaluating Relevancy
Relevancy evaluation is the process of continuously verifying that the information retrieved and generated by the agent is appropriate and accurate for the user's query.
Core concepts of relevancy evaluation:
| Concept | Description |
|---|---|
| Context Awareness | Understanding query context to retrieve relevant information (e.g., considering user budget and preferences) |
| Accuracy | Provided information must be factually correct and up-to-date |
| User Intent | Inferring the purpose behind the query to provide the most relevant information |
| Feedback Loop | Continuously collecting and analyzing user feedback to refine the relevancy evaluation process |
Practical relevancy evaluation techniques:
relevancy_techniques.pypython
Complete example applying relevancy evaluation to a Travel Agent:
travel_agent_relevancy.pypython
Relevancy Evaluation FlowMermaidflowchart TD Pref["User Preferences\n{destination, budget, interests}"] --> GenRec["generate_recommendations()\nStart itinerary generation"] GenRec --> Retrieve["retrieve_information()\nSearch flights, hotels, attractions"] Retrieve --> Rank["filter_and_rank(hotels)\nSort by relevance_score → Top 10"] Rank --> Score["relevance_score(item, query)\nScore matching category, price, location"] Score --> Itin["create_itinerary()\nGenerate final itinerary"] Itin --> FB["User Feedback\n{liked, disliked}"] FB --> Adjust["adjust_based_on_feedback()\nAdjust relevance +1 / -1"] Adjust -->|"Reflect in next recommendation"| Rank
Search with Intent
Intent-aware Search goes beyond keyword matching to identify the actual intent behind a user's query, providing more relevant results.
Three intent types:
| Intent | Description | Example |
|---|---|---|
| Informational | Information gathering | "What are good museums to visit in Paris?" |
| Navigational | Navigating to a specific website or page | "Official Louvre Museum website" |
| Transactional | Executing transactions such as booking or purchasing | "Book flight to Paris" |
search_with_intent.pypython
Intent-Aware Search FlowMermaidzenuml title Intent-Aware Search Flow User->Agent: "best museums in Paris" Agent->IntentClassifier: identify_intent(query) IntentClassifier->Agent: "informational" Agent->ContextAnalyzer: analyze_context(query, user_history) ContextAnalyzer->Agent: current + history context Agent->SearchEngine: search informational query with preferences SearchEngine->Agent: raw results Agent->Personalizer: personalize_results(results, user_history) Personalizer->Agent: filtered top 10 Agent->User: personalized recommendations
Generating Code as a Tool
Code Generating Agents use AI models to write and execute code to solve complex problems.
Key application areas: automated code generation, SQL as RAG, and automated data analysis.
Key application areas: automated code generation, SQL as RAG, and automated data analysis.
Implemented step by step using a Travel Agent as an example.
Step 1 — Gather User Preferences
step1_agent.pypython
Step 2 — Dynamically Generate Data Retrieval Code
The agent directly generates code snippets tailored to preferences.
step2_generate_code.pypython
Step 3 — Execute Generated Code
step3_execute.pypython
Step 4 — Generate Itinerary
step4_itinerary.pypython
Step 5 — Regenerate Code Based on Feedback
Update preferences reflecting feedback, then regenerate and re-execute code with the updated preferences.
step5_feedback_regen.pypython
Code Generating Agent FlowMermaidflowchart TD S1["Step 1\ngather_preferences()\nCollect user preferences"] --> S2["Step 2\ngenerate_code_to_fetch_data()\ngenerate_code_to_fetch_hotels()\nDynamically generate code"] --> S3["Step 3\nexecute_code()\nExecute generated code → flights, hotels"] --> S4["Step 4\ngenerate_itinerary()\nCombine flights, hotels, attractions"] --> S5["Step 5\nadjust_based_on_feedback()\nUpdate preferences → Regenerate & re-execute code"] S5 -->|"Updated Preferences"| S2
Leveraging Environmental Awareness and Reasoning
Schema-based environmental awareness is an approach where the agent understands table schemas and reasons about which fields to adjust and how based on feedback.
schema_awareness.pypython
Schema-Aware Feedback Adjustment FlowMermaidflowchart TD FB["User Feedback\n{liked, disliked}"] SC["Schema\n{favorites·avoid: +/-/default}"] FB & SC --> AdjFn["adjust_based_on_feedback()\nDirectly update favorites & avoid"] AdjFn --> Loop["Iterate over schema fields"] Loop --> EnvFn["adjust_based_on_environment()\nDetermine adjustment value per field"] EnvFn --> UPref["Updated Preferences"] UPref --> GenF["generate_code_to_fetch_data()"] UPref --> GenH["generate_code_to_fetch_hotels()"] GenF --> ExF["execute_code() → flights"] GenH --> ExH["execute_code() → hotels"] ExF & ExH --> Itin["generate_itinerary()\nUpdated Itinerary"]
Using SQL as a Retrieval-Augmented Generation (RAG) Technique
SQL as RAG generates dynamic SQL queries based on user input to retrieve information from a database:
sql_as_rag.pypython
Example dynamically generated SQL queries:
example_queries.sqlsql
SQL as RAG: Query Generation & Execution FlowMermaidflowchart TD Pref["User Preferences\n{destination, budget, dates, interests}"] --> GenRec["generate_recommendations(preferences)"] GenRec --> FQ["generate_sql_query('flights')\nSELECT * FROM flights WHERE ..."] GenRec --> HQ["generate_sql_query('hotels')\nSELECT * FROM hotels WHERE ..."] GenRec --> AQ["generate_sql_query('attractions')\nSELECT * FROM attractions WHERE ..."] FQ --> EF["execute_sql_query()\n→ flight results"] HQ --> EH["execute_sql_query()\n→ hotel results"] AQ --> EA["execute_sql_query()\n→ attraction results"] EF & EH & EA --> Dict["{flights, hotels, attractions}\nItinerary Dict"]
Example of Metacognition
Here is a metacognition implementation example that integrates the concepts learned so far.
It implements a 3-step cycle where the agent makes an initial decision → recognizes errors through self-reflection → adjusts its strategy:
It implements a 3-step cycle where the agent makes an initial decision → recognizes errors through self-reflection → adjusts its strategy:
hotel_metacognition.pypython
Metacognition: Reflect & Adjust LoopMermaidflowchart TD Start["Initial Decision\nrecommend_hotel('cheapest')"] --> Eval["Self-Reflection\nreflect_on_choice()"] Eval --> Feedback{User Feedback} Feedback -->|"'bad'\nprice < 100 or quality < 7"| Switch["Switch Strategy\n'cheapest' → 'highest_quality'"] Switch --> Retry["Re-recommend\nrecommend_hotel('highest_quality')"] Retry --> Eval Feedback -->|"'good'"| Done["Confirm Recommendation"]
This is the essence of metacognition:
- Initial Decision: Selected Budget Inn with the 'cheapest' strategy
- Self-Reflection: Recognized "bad" feedback due to quality 6
- Strategy Adjustment: Ultimately recommended Luxury Stay
The agent does not simply alter the final recommendation; it modifies its own decision-making mechanism.
Summary
Lesson 09 SummaryMermaidflowchart LR Root["Metacognition\nin AI Agents"] Root --> MC["Metacognition Loop"] Root --> RAG["RAG Strategies"] Root --> Search["Intent Search"] Root --> Code["Code Generation"] MC --> M1["Self-Reflection\nPerformance Evaluation"] MC --> M2["Error Detection\nError Recognition"] MC --> M3["Strategy Adjustment\nModifying Strategy"] RAG --> R1["Corrective RAG\nFeedback-based Re-search"] RAG --> R2["Pre-emptive Load\nPre-load Context"] RAG --> R3["SQL as RAG\nDynamic Query Generation"] Search --> S1["Informational"] Search --> S2["Navigational"] Search --> S3["Transactional"] Code --> C1["generate_code\nDynamic Code Generation"] Code --> C2["execute_code\nExecution & Results Reflection"]
- Metacognition is the ability of an agent to recognize and modify its own decision-making process, representing an improvement of the reasoning strategy itself rather than a simple correction of outputs.
- Corrective RAG implements an iterative self-correction loop that incorporates user feedback into preferences/avoids, rewrites queries, and performs re-retrieval.
- Pre-emptive Context Load pre-loads known domain knowledge into memory to provide instant responses without external retrieval, making it highly effective when latency is critical.
- Intent-aware Search classifies queries into Informational, Navigational, or Transactional intents, delivering far more relevant results than simple keyword matching.
- Code Generation + SQL as RAG empowers agents to dynamically write and execute code to retrieve real-time data and automate complex tasks.