AI Agents for Beginners - 9. Metacognition in AI Agents

Metacognition in AI agents — self-reflection, corrective RAG, pre-emptive context loading, goal bootstrapping, intent-aware search, code generation, and SQL as RAG.
June 9, 2026

AI Agents for Beginners - 9. Metacognition in AI Agents

This post summarizes Lesson 09 from Microsoft's AI Agents for Beginners course.

What is Metacognition?

Metacognition is "thinking about thinking."
In AI agents, metacognition refers to the ability of an agent to recognize its own internal processes and monitor, regulate, and adapt its behavior.
Key challenges addressed by metacognition:
ChallengeDescription
TransparencyThe agent must be able to explain its reasoning and decisions
ReasoningEnhances the ability to synthesize information and make sound decisions
AdaptationAdapts to new environments and changing conditions
PerceptionImproves accuracy in perceiving and interpreting environmental data
Why metacognition is important in agent design:
Importance of Metacognition
Mermaid
flowchart LR MC["Metacognition\nRecognizing Own Thought Process"] MC --> SR["Self-Reflection\nEvaluating Performance & Identifying Improvement Areas"] MC --> AD["Adaptability\nAdjusting Strategy Based on Past Experience"] MC --> EC["Error Correction\nAutonomous Error Detection & Correction"] MC --> RM["Resource Management\nOptimizing Time & Compute Resources"]
Core examples of metacognition:
  • "I prioritized cheap flights... but I might have missed direct flights. Let me check again."
  • "Whenever the user said 'it's too complicated,' the issue was that I recommended items based on popularity."
The agent components—Persona, Tools, and Skills—combine to form an "expertise unit" that executes specific tasks, and metacognition enables this unit to improve on its own.

Planning in Agents

Planning is a core element of AI agent behavior that designs steps to achieve a goal while taking into account the current state, resources, and obstacles.
9-step planning process of a travel agent:
Travel Agent Planning Flow
Mermaid
flowchart TD P1["1. Gather Preferences\nCollect dates, budget, interests"] --> P2["2. Retrieve Information\nSearch flights, hotels, attractions"] --> P3["3. Generate Recommendations\nCreate personalized itinerary"] --> P4["4. Present to User\nShare itinerary"] --> P5["5. Collect Feedback\nGather user feedback"] --> P6["6. Adjust Based on Feedback\nUpdate itinerary"] --> P7["7. Final Confirmation\nFinal review"] --> P8["8. Book & Confirm\nProceed with booking"] --> P9["9. Ongoing Support\nSupport during trip"]
Travel Agent implementation applying metacognition:
travel_agent_metacognition.py
python
Metacognition Loop
Mermaid
flowchart TD G["gather_preferences()\nInitialize preferences"] --> R["generate_recommendations()\nSearch flights/hotels/attractions + Generate itinerary"] --> U["Present itinerary to user"] --> FB["User Feedback\n{liked, disliked}"] FB --> Adj["adjust_based_on_feedback()\nAccumulate experience_data\nReadjust preferences"] Adj -->|"Updated Preferences"| R
Accumulating feedback in experience_data, readjusting preferences with adjust_based_on_feedback, and generating a new itinerary forms the core of the metacognition loop.

Corrective RAG System

Comparing two information access strategies:
Corrective RAGPre-emptive Context Load
ApproachRetrieve information from external sources at query timePre-load relevant context prior to processing
UpdatesReal-time informationPre-loaded information
SpeedIncurs retrieval latencyImmediate response
AccuracyFeedback-based error correctionFixed context
Best ForDynamic data, when error correction is neededKnown domain, when fast responses are required
RAG Tool Calling vs Pre-Emptive Context Load
Mermaid
%%{init: {'look': 'handDrawn'}}%% flowchart TD subgraph RAG["RAG Tool Calling"] direction TD LLM1["LLM"] LLM1 --> T1["Tool"] LLM1 --> T2["Tool"] LLM1 --> T3["Tool"] T1 & T2 & T3 --> R1["Response"] end subgraph PCL["Pre-Emptive Context Load"] direction TD LLM2["LLM"] LLM2 --> CTX["Context\n(Provided before query processing)"] CTX --> R2["Response"] end

Corrective RAG Approach

Corrective RAG corrects errors and improves accuracy through three components:
ElementDescription
Prompting TechniqueCrafting specific prompts that guide the agent to retrieve relevant information
ToolImplementing algorithms and mechanisms to evaluate the relevance of retrieved information and generate accurate responses
EvaluationContinuously evaluating agent performance and iterating on adjustments to improve accuracy and efficiency

Step-by-Step Implementation of Corrective RAG

Five steps to applying Corrective RAG to a Travel Agent:
Step 1 — Gather User Input
Collect travel preferences such as destination, dates, budget, and interests.
step1_preferences.py
python
Step 2 — Retrieve Information and Generate Initial Recommendations
Search for flights, hotels, and attractions based on the collected preferences and generate an initial itinerary.
step2_retrieve.py
python
Step 3 — Collect User Feedback
Collect positive and negative user feedback on the initial recommendations.
step3_feedback.py
python
Step 4 — Apply Corrective RAG
  • Prompting Technique: Formulate a new search query reflecting the feedback
step4_prompting.py
python
  • Tool: Re-search with updated preferences and generate a new itinerary
step4_tool.py
python
  • Evaluation: Continuous evaluation function reflecting feedback into preferences
step4_evaluation.py
python
Step 5 — Putting It Together: Corrective RAG Travel Agent
This final example integrates the above steps into a single class. adjust_based_on_feedback handles feedback accumulation → preference update → new itinerary generation all in one go.
corrective_rag_travel_agent.py
python
Corrective RAG Travel Agent Flow
Mermaid
flowchart TD GP["gather_preferences()\nStore destination, budget, interests"] --> GR["generate_recommendations()\nCall retrieve_information()"] GR --> RT["search_flights + hotels + attractions"] RT --> CI["create_itinerary()"] CI --> SU["Present itinerary to user"] SU --> FB["User Feedback\n{liked, disliked}"] FB --> AF["adjust_based_on_feedback()\nexperience_data.append(feedback)\nadjust_preferences() → Update preferences"] AF -->|"Updated user_preferences"| GR

Corrective RAG Flow

Corrective RAG Flow
Mermaid
flowchart TD User["User Input\nDestination, budget, interests"] --> Retrieve["Retrieve Information\nFlights, hotels, attractions"] --> Gen["Generate Initial Itinerary"] --> Feedback["Collect User Feedback"] Feedback --> Adjust{Analyze Feedback} Adjust -->|"disliked items"| NewQuery["Rewrite Query\nAdd preferences['avoid']"] NewQuery --> Retrieve Adjust -->|"Satisfied"| Done["Finalize Itinerary"]

Pre-emptive Context Load

Pre-loads known domain knowledge into memory before processing to respond immediately without external retrieval:
preemptive_context.py
python
Pre-emptive Context Load Flow
Mermaid
flowchart TD Init["__init__()\nPre-load self.context\n{Paris, Tokyo, NY, Sydney}"] --> Query["User destination query"] --> Get["get_destination_info(destination)\nself.context.get(destination)"] Get --> Check{Context exists?} Check -->|"Yes"| Hit["country, currency, language, attractions\nReturn immediately (no external retrieval)"] Check -->|"No"| Miss["Return 'No information available'"]

Bootstrapping a Plan with a Goal Before Iterating

An approach where a goal is defined upfront, followed by iterative refinement of the plan.
Initial planning is established with bootstrap_plan, and progressively optimized using iterate_plan:
bootstrap_plan.py
python
Goal Bootstrapping Loop
Mermaid
flowchart TD Goal["Define Clear Goal\n(e.g., maximize satisfaction, budget ≤ $2000)"] --> Bootstrap["bootstrap_plan\nCreate initial plan based on goal & budget"] --> Execute["Execute Plan"] --> Eval{Evaluate Results} Eval -->|"Can be improved"| Iterate["iterate_plan\nReplace with better destinations"] Iterate --> Execute Eval -->|"Optimal"| Done["Finalize Travel Plan"]

Taking advantage of LLM for Re-ranking and Scoring

LLM Re-ranking is a technique where instead of using initial retrieval results directly, an LLM evaluates the relevance and quality of each candidate and reorders them into an optimal sequence.
StageDescription
RetrievalInitial retrieval of candidate documents/responses based on the query
Re-rankingLLM reorders candidates by relevance and quality — placing the most helpful information at the top
ScoringLLM assigns scores to each candidate to select the optimal response or document
Example of an LLM re-ranking and scoring destinations based on user preferences in a travel agent:
llm_reranking.py
python
LLM Re-ranking & Scoring Flow
Mermaid
flowchart TD Query["User Preferences\nactivity: sightseeing\nculture: diverse"] --> Retrieve["Initial Candidate List\n(Paris, Tokyo, NYC, Sydney)"] --> Prompt["generate_prompt()\nInclude preferences + destination descriptions"] --> LLM["LLM (Azure OpenAI)\nEvaluate relevance & quality"] LLM --> Rerank["Re-ranking\nSort best-matching destinations to the top"] Rerank --> Score["Scoring\nAssign scores to each candidate"] Score --> Result["Final Recommendations\n(Including ranks & scores)"]

RAG: Prompting Technique vs Tool

RAG can be utilized in two ways: as a prompting technique or as an integrated tool:
AspectPrompting TechniqueTool
ApproachManually crafting prompts for each queryAutomating retrieval and generation
ControlFine-grained control over the retrieval processFull end-to-end automated handling
FlexibilityCustom prompts tailored to specific requirementsMore efficient for large-scale implementations
ComplexityRequires writing and tuning promptsEasy to integrate into AI Agent architectures
rag_prompting_vs_tool.py
python

Evaluating Relevancy

Relevancy evaluation is the process of continuously verifying that the information retrieved and generated by the agent is appropriate and accurate for the user's query.
Core concepts of relevancy evaluation:
ConceptDescription
Context AwarenessUnderstanding query context to retrieve relevant information (e.g., considering user budget and preferences)
AccuracyProvided information must be factually correct and up-to-date
User IntentInferring the purpose behind the query to provide the most relevant information
Feedback LoopContinuously collecting and analyzing user feedback to refine the relevancy evaluation process
Practical relevancy evaluation techniques:
relevancy_techniques.py
python
Complete example applying relevancy evaluation to a Travel Agent:
travel_agent_relevancy.py
python
Relevancy Evaluation Flow
Mermaid
flowchart TD Pref["User Preferences\n{destination, budget, interests}"] --> GenRec["generate_recommendations()\nStart itinerary generation"] GenRec --> Retrieve["retrieve_information()\nSearch flights, hotels, attractions"] Retrieve --> Rank["filter_and_rank(hotels)\nSort by relevance_score → Top 10"] Rank --> Score["relevance_score(item, query)\nScore matching category, price, location"] Score --> Itin["create_itinerary()\nGenerate final itinerary"] Itin --> FB["User Feedback\n{liked, disliked}"] FB --> Adjust["adjust_based_on_feedback()\nAdjust relevance +1 / -1"] Adjust -->|"Reflect in next recommendation"| Rank

Search with Intent

Intent-aware Search goes beyond keyword matching to identify the actual intent behind a user's query, providing more relevant results.
Three intent types:
IntentDescriptionExample
InformationalInformation gathering"What are good museums to visit in Paris?"
NavigationalNavigating to a specific website or page"Official Louvre Museum website"
TransactionalExecuting transactions such as booking or purchasing"Book flight to Paris"
search_with_intent.py
python
Intent-Aware Search Flow
Mermaid
zenuml title Intent-Aware Search Flow User->Agent: "best museums in Paris" Agent->IntentClassifier: identify_intent(query) IntentClassifier->Agent: "informational" Agent->ContextAnalyzer: analyze_context(query, user_history) ContextAnalyzer->Agent: current + history context Agent->SearchEngine: search informational query with preferences SearchEngine->Agent: raw results Agent->Personalizer: personalize_results(results, user_history) Personalizer->Agent: filtered top 10 Agent->User: personalized recommendations

Generating Code as a Tool

Code Generating Agents use AI models to write and execute code to solve complex problems.
Key application areas: automated code generation, SQL as RAG, and automated data analysis.
Implemented step by step using a Travel Agent as an example.
Step 1 — Gather User Preferences
step1_agent.py
python
Step 2 — Dynamically Generate Data Retrieval Code
The agent directly generates code snippets tailored to preferences.
step2_generate_code.py
python
Step 3 — Execute Generated Code
step3_execute.py
python
Step 4 — Generate Itinerary
step4_itinerary.py
python
Step 5 — Regenerate Code Based on Feedback
Update preferences reflecting feedback, then regenerate and re-execute code with the updated preferences.
step5_feedback_regen.py
python
Code Generating Agent Flow
Mermaid
flowchart TD S1["Step 1\ngather_preferences()\nCollect user preferences"] --> S2["Step 2\ngenerate_code_to_fetch_data()\ngenerate_code_to_fetch_hotels()\nDynamically generate code"] --> S3["Step 3\nexecute_code()\nExecute generated code → flights, hotels"] --> S4["Step 4\ngenerate_itinerary()\nCombine flights, hotels, attractions"] --> S5["Step 5\nadjust_based_on_feedback()\nUpdate preferences → Regenerate & re-execute code"] S5 -->|"Updated Preferences"| S2

Leveraging Environmental Awareness and Reasoning

Schema-based environmental awareness is an approach where the agent understands table schemas and reasons about which fields to adjust and how based on feedback.
schema_awareness.py
python
Schema-Aware Feedback Adjustment Flow
Mermaid
flowchart TD FB["User Feedback\n{liked, disliked}"] SC["Schema\n{favorites·avoid: +/-/default}"] FB & SC --> AdjFn["adjust_based_on_feedback()\nDirectly update favorites & avoid"] AdjFn --> Loop["Iterate over schema fields"] Loop --> EnvFn["adjust_based_on_environment()\nDetermine adjustment value per field"] EnvFn --> UPref["Updated Preferences"] UPref --> GenF["generate_code_to_fetch_data()"] UPref --> GenH["generate_code_to_fetch_hotels()"] GenF --> ExF["execute_code() → flights"] GenH --> ExH["execute_code() → hotels"] ExF & ExH --> Itin["generate_itinerary()\nUpdated Itinerary"]

Using SQL as a Retrieval-Augmented Generation (RAG) Technique

SQL as RAG generates dynamic SQL queries based on user input to retrieve information from a database:
sql_as_rag.py
python
Example dynamically generated SQL queries:
example_queries.sql
sql
SQL as RAG: Query Generation & Execution Flow
Mermaid
flowchart TD Pref["User Preferences\n{destination, budget, dates, interests}"] --> GenRec["generate_recommendations(preferences)"] GenRec --> FQ["generate_sql_query('flights')\nSELECT * FROM flights WHERE ..."] GenRec --> HQ["generate_sql_query('hotels')\nSELECT * FROM hotels WHERE ..."] GenRec --> AQ["generate_sql_query('attractions')\nSELECT * FROM attractions WHERE ..."] FQ --> EF["execute_sql_query()\n→ flight results"] HQ --> EH["execute_sql_query()\n→ hotel results"] AQ --> EA["execute_sql_query()\n→ attraction results"] EF & EH & EA --> Dict["{flights, hotels, attractions}\nItinerary Dict"]

Example of Metacognition

Here is a metacognition implementation example that integrates the concepts learned so far.
It implements a 3-step cycle where the agent makes an initial decision → recognizes errors through self-reflection → adjusts its strategy:
hotel_metacognition.py
python
Metacognition: Reflect & Adjust Loop
Mermaid
flowchart TD Start["Initial Decision\nrecommend_hotel('cheapest')"] --> Eval["Self-Reflection\nreflect_on_choice()"] Eval --> Feedback{User Feedback} Feedback -->|"'bad'\nprice < 100 or quality < 7"| Switch["Switch Strategy\n'cheapest' → 'highest_quality'"] Switch --> Retry["Re-recommend\nrecommend_hotel('highest_quality')"] Retry --> Eval Feedback -->|"'good'"| Done["Confirm Recommendation"]
This is the essence of metacognition:
  • Initial Decision: Selected Budget Inn with the 'cheapest' strategy
  • Self-Reflection: Recognized "bad" feedback due to quality 6
  • Strategy Adjustment: Ultimately recommended Luxury Stay
The agent does not simply alter the final recommendation; it modifies its own decision-making mechanism.

Summary

Lesson 09 Summary
Mermaid
flowchart LR Root["Metacognition\nin AI Agents"] Root --> MC["Metacognition Loop"] Root --> RAG["RAG Strategies"] Root --> Search["Intent Search"] Root --> Code["Code Generation"] MC --> M1["Self-Reflection\nPerformance Evaluation"] MC --> M2["Error Detection\nError Recognition"] MC --> M3["Strategy Adjustment\nModifying Strategy"] RAG --> R1["Corrective RAG\nFeedback-based Re-search"] RAG --> R2["Pre-emptive Load\nPre-load Context"] RAG --> R3["SQL as RAG\nDynamic Query Generation"] Search --> S1["Informational"] Search --> S2["Navigational"] Search --> S3["Transactional"] Code --> C1["generate_code\nDynamic Code Generation"] Code --> C2["execute_code\nExecution & Results Reflection"]
  • Metacognition is the ability of an agent to recognize and modify its own decision-making process, representing an improvement of the reasoning strategy itself rather than a simple correction of outputs.
  • Corrective RAG implements an iterative self-correction loop that incorporates user feedback into preferences/avoids, rewrites queries, and performs re-retrieval.
  • Pre-emptive Context Load pre-loads known domain knowledge into memory to provide instant responses without external retrieval, making it highly effective when latency is critical.
  • Intent-aware Search classifies queries into Informational, Navigational, or Transactional intents, delivering far more relevant results than simple keyword matching.
  • Code Generation + SQL as RAG empowers agents to dynamically write and execute code to retrieve real-time data and automate complex tasks.

Jooojub
System S/W engineer
Explore Tags
Series
    Recent Post
    © 2026. jooojub. All right reserved.