AI Agents for Beginners - 14. Microsoft Agent Framework
Microsoft Agent Framework (MAF) core concepts — Agents, Threads, Middleware, Memory, Observability — and advanced workflow patterns including Sequential, Concurrent, Handoff orchestration, Middleware injection, and Human-in-the-Loop with code examples.
June 9, 2026
AI Agents for Beginners - 14. Microsoft Agent Framework
This post is a summary based on Lesson 14 of Microsoft's AI Agents for Beginners course.
Introducing Microsoft Agent Framework (MAF), which allows you to implement a wide range of AI Agent concepts at a production level.
Understanding Microsoft Agent Framework
Microsoft Agent Framework (MAF) is an integrated framework that provides the flexibility to handle diverse agent use cases in both production and research environments.
Orchestration Types
MAF supports 5 orchestration patterns depending on how agents collaborate:
| Orchestration | Description |
|---|---|
| Sequential | Sequential agent orchestration in scenarios requiring step-by-step workflows |
| Concurrent | Concurrent orchestration in scenarios where agents need to complete tasks simultaneously |
| Group Chat | Group chat orchestration in scenarios where agents collaborate together on a single task |
| Handoff | Handoff orchestration where agents pass tasks to each other as subtasks are completed |
| Magnetic | Magnetic orchestration where a manager agent creates a task list and coordinates sub-agents |
SequentialMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A1[Agent 1] --> A2[Agent 2] --> A3[Agent 3]
ConcurrentMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR D[Dispatcher] --> A1[Agent 1] & A2[Agent 2] & A3[Agent 3]
Group ChatMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent A] <--> B[Agent B] B <--> C[Agent C] A <--> C
HandoffMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A1[Agent 1] -->|subtask done| A2[Agent 2] -->|subtask done| A3[Agent 3]
MagneticMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD M[Manager] --> TL[Task List] TL --> S1[Sub-Agent 1] & S2[Sub-Agent 2] & S3[Sub-Agent 3]
Production Features
MAF goes beyond simple agent execution by embedding enterprise-grade capabilities:
| Feature | Description |
|---|---|
| Observability | Track all actions (tool calls, reasoning flows) via OpenTelemetry and monitor performance with the Microsoft Foundry dashboard |
| Security | Includes role-based access control, personal data handling, and built-in safety content |
| Durability | Long-running processes where agent threads and workflows can be paused, resumed, and recovered from errors |
| Control | Supports Human-in-the-Loop workflows for tasks requiring human approval |
ObservabilityMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent] --> T[Tool Call] --> O[OTel Trace] --> D[Dashboard]
SecurityMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR R[Request] --> RB[RBAC Check] --> A[Agent]
DurabilityMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR R[Run] --> P[Pause] --> RS[Resume] --> C[Complete]
ControlMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent] --> H[Human Approval] --> C[Continue]
Interoperability
MAF is not locked to a specific vendor:
| Feature | Description |
|---|---|
| Cloud-agnostic | Run agents across containers, on-premises, and multi-cloud environments |
| Provider-agnostic | Create agents using your preferred SDK, including Azure OpenAI, OpenAI, etc. |
| Open Standards | Leverage protocols such as A2A (Agent-to-Agent) and MCP (Model Context Protocol) |
| Plugins and Connectors | Connect to data services like Microsoft Fabric, SharePoint, Pinecone, and Qdrant |
Cloud-agnosticMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent] --> C1[Container] & C2[On-prem] & C3[Multi-Cloud]
Provider-agnosticMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent] --> P1[Azure OpenAI] & P2[OpenAI] & P3[MiniMax]
Open StandardsMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent] <--> S1[A2A] & S2[MCP]
Plugins & ConnectorsMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Agent] --> P1[Fabric] & P2[SharePoint] & P3[Pinecone/Qdrant]
Key Concepts of Microsoft Agent Framework
Agents
Creating Agents
Agents are created by defining an inference service (LLM provider), a set of instructions, and a
name:A variety of LLM providers are supported — Azure AI Foundry Agent Service:
OpenAI or OpenAI-compatible APIs like MiniMax, which offers a 204K token context window, can also be used:
Remote agents using the A2A protocol are also supported:
Running Agents
Agents can be executed with non-streaming or streaming responses:
Non-Streaming (agent.run)Mermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant User participant Agent participant LLM User->>Agent: run("question") Agent->>LLM: request Note over LLM: generates full response LLM-->>Agent: complete response Agent-->>User: result.text (all at once)
Streaming (agent.run_stream)Mermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant User participant Agent participant LLM User->>Agent: run_stream("question") Agent->>LLM: request loop token by token LLM-->>Agent: chunk Agent-->>User: update.text (partial) end
In each agent execution, parameters such as
max_tokens, tools, and model can be customized.Tools
Tools can be provided when defining the agent or at runtime:
Agent Threads
Agent threads handle multi-turn conversations. They can be created with
get_new_thread() or automatically generated upon execution:Threads can be serialized for storage and reused in subsequent sessions:
Sequence DiagramMermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant User participant Agent participant Thread User->>Agent: run("Where to go?", thread=thread) Agent->>Thread: append message Agent-->>User: response User->>Agent: run("What about hotels?", thread=thread) Agent->>Thread: append message + read history Agent-->>User: response (with context) Agent->>Thread: serialize() Thread-->>Agent: serialized_thread Note over Thread: Can be reused in subsequent sessions
Agent Middleware
Use middleware to execute tasks or trace actions between the agent, tools, and LLMs.
Function Middleware — executes between the agent and function/tool calls:
Chat Middleware — executes between the agent and AI service requests:
Both middlewares insert pre-processing/post-processing logic before and after calling
next(context):Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR subgraph FunctionMiddleware["Function Middleware"] direction LR Pre1["Pre-processing\n(Logging, etc.)"] --> Func["Actual Function\nExecution"] --> Post1["Post-processing\n(Inspect results)"] end subgraph ChatMiddleware["Chat Middleware"] direction LR Pre2["Pre-processing\n(Check message count)"] --> LLM["LLM\nCall"] --> Post2["Post-processing\n(Log response)"] end Agent --> FunctionMiddleware Agent --> ChatMiddleware
Agent Memory
MAF provides 3 types of memory:
In-Memory Storage — retained only during application runtime:
In-Memory StorageMermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant User participant Agent participant Thread User->>Agent: run(message, thread=thread) Agent->>Thread: store message Agent-->>User: response Note over Thread: Cleared when app terminates
Persistent Messages — stores conversation history across sessions:
Persistent MessagesMermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant User participant Agent participant ChatMessageStore User->>Agent: Session 1 Agent->>ChatMessageStore: save messages Note over Agent: session ends User->>Agent: Session 2 Agent->>ChatMessageStore: load history ChatMessageStore-->>Agent: previous messages Agent-->>User: response (with context)
Dynamic Memory — added to context before agents run, stored in external services like Mem0:
Dynamic MemoryMermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant User participant Agent participant Mem0 participant LLM User->>Agent: run(message) Agent->>Mem0: fetch memories Mem0-->>Agent: relevant context Agent->>LLM: messages + memory context LLM-->>Agent: response Agent-->>User: response
Agent Observability
MAF integrates with OpenTelemetry to provide tracing and metrics:
Workflows
MAF provides workflows to complete tasks through predefined steps. It supports multi-agent orchestration and Workflow Checkpointing.
Executors
Executors receive input messages, perform assigned tasks, and generate output messages. Both AI agents and custom logic can serve as executors.
Edges
Edges define the flow of messages within a workflow:
| Edge Type | Description |
|---|---|
| Direct Edge | Simple 1 connection between executors |
| Conditional Edge | Activated after a certain condition is met |
| Switch-case Edge | Routes messages to different executors based on defined conditions |
| Fan-out Edge | Sends one message to multiple targets |
| Fan-in Edge | Collects messages from multiple executors and sends to one target |
Direct Edge Example:
Direct EdgeMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Executor A] --> B[Executor B]
Conditional EdgeMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Executor A] -->|"condition met"| B[Executor B] A -->|"condition not met"| A
Switch-case EdgeMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Executor A] -->|"case 1"| B[Executor B] A -->|"case 2"| C[Executor C] A -->|"case 3"| D[Executor D]
Fan-out EdgeMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Executor A] --> B[Executor B] A --> C[Executor C] A --> D[Executor D]
Fan-in EdgeMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR A[Executor A] --> D[Executor D] B[Executor B] --> D C[Executor C] --> D
Events
For workflow observability, MAF provides the following execution events:
| Event | Triggered When |
|---|---|
WorkflowStartedEvent | Workflow execution starts |
WorkflowOutputEvent | Workflow output is generated |
WorkflowErrorEvent | An error occurs in the workflow |
ExecutorInvokeEvent | Executor task starts |
ExecutorCompleteEvent | Executor task completes |
RequestInfoEvent | A request is published |
Workflow Events — Hooking PointsMermaid%%{init: {'look': 'handDrawn'}}%% sequenceDiagram participant App participant Workflow participant Executor participant Logger as Logger / Hook App->>Workflow: workflow.run_stream(input) Workflow-->>Logger: WorkflowStartedEvent loop Each Executor in sequence Workflow->>Executor: invoke Executor-->>Logger: ExecutorInvokeEvent Note over Executor: Process task Executor-->>Logger: ExecutorCompleteEvent Executor->>Workflow: output end alt Human input required Workflow-->>Logger: RequestInfoEvent Logger->>App: Prompt user with question App->>Workflow: send_responses_streaming() else Error occurs Workflow-->>Logger: WorkflowErrorEvent end Workflow-->>Logger: WorkflowOutputEvent Workflow-->>App: Final result
Advanced MAF Patterns
Sequential Orchestration
A pattern where two agents work sequentially. The Front Desk Agent provides initial recommendations, and the Concierge Agent reviews and evaluates them:
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR U["User Input"] --> FD["Front Desk Agent\nProvide Recommendations"] FD --> C["Concierge Agent\nReview & Evaluate"] C --> O["Final Output\nRecommendation + Expert Review"]
Key benefits of sequential orchestration:
- Iterative Refinement: The second agent refines the work of the first agent
- Specialization: Each agent handles a specific role in the process
- Quality Control: Built-in review and verification steps
Concurrent Orchestration
A pattern where multiple agents are executed concurrently in parallel. Uses the Fan-out/Fan-in pattern:
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR U["User Input\nTokyo travel recommendations"] --> D["InputDispatcher\nFan-out broadcast"] D --> A["Attractions Agent\nAttractions & Activities"] D --> B["Dining Agent\nFood & Restaurants"] D --> C["History Agent\nHistory & Culture"] A --> O["Integrated Output\nComprehensive Travel Guide"] B --> O C --> O
Performance: Compared to sequential execution, concurrent execution is much faster. Running three agents sequentially requires summing up each response time, whereas concurrent execution takes only as long as the slowest agent.
Handoff Orchestration
A pattern that dynamically delegates tasks to specialist agents depending on the type of customer request:
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD U["User Request"] --> CS["Customer Support Agent\n(Coordinator)"] CS -->|"Flight booking request"| BA["Booking Agent\n✈️ Book flight"] CS -->|"Refund/dispute request"| DA["Disputes Agent\n💰 Process refund"] CS -->|"Trip confirmation request"| TA["Trip Check Agent\n🎯 Confirm itinerary"] BA --> R["Resolution\nSpecialist handling complete"] DA --> R TA --> R
Core benefits of handoff orchestration:
- Dynamic Routing: The Customer Support agent analyzes request intent to route to specialists
- Context Preservation: Retains full conversation history across transitions (no need for users to repeat information)
- Specialization: Each agent focuses solely on its domain of expertise
Middleware Injection
A pattern that uses middleware to intercept and modify results of an agent's tool function executions.
Example: Priority members are automatically overridden to "Available" even if there is no hotel availability:
Middleware is injected via the
middleware parameter during agent creation:Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart LR subgraph Regular["Regular User"] direction LR R1["Paris Request"] --> R2["hotel_booking\nParis unavailable"] --> R3["alternative_agent\nSuggest alternatives"] end subgraph Priority["Priority Member"] direction LR P1["Paris Request"] --> P2["hotel_booking\nParis unavailable"] --> MW["🌟 Middleware\nOverride!"] --> P3["booking_agent\nProceed with booking"] end
Decorator vs Middleware:
| Feature | Decorator | Middleware |
|---|---|---|
| Scope | Single function | All functions in agent |
| Flexibility | Fixed at definition | Dynamic at runtime |
| Context | Limited | Full agent context |
| Agent-Aware | No | Yes |
Human-in-the-Loop
A pattern where an AI agent pauses execution and requests human input before proceeding. Essential for critical decisions, ambiguous situations, and scenarios requiring compliance.
MAF provides 3 core components:
RequestInfoExecutor: Pauses the workflow and emitsRequestInfoEventRequestInfoMessage: Base class for request payloads sent to humansRequestResponse: Links requests and responses viarequest_id
Workflow configuration:
Human-in-the-Loop execution is handled in an event-driven manner:
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD Start["availability_agent\nCheck availability"] -->|"No availability"| Conf["confirmation_agent\nGenerate confirmation question"] Start -->|"Available"| Book["booking_agent\nRecommend booking"] Conf --> Prep["prepare_human_request\nType conversion"] Prep --> RI["RequestInfoExecutor\n⏸️ PAUSE"] RI --> DM["DecisionManager\nProcess human response"] DM -->|"yes"| Alt["alternative_agent\nSuggest alternatives"] DM -->|"no"| Cancel["cancellation_agent\nProcess cancellation"] Alt --> End["display_result"] Cancel --> End Book --> End
Advanced MAF Patterns Summary
Advanced patterns to consider when building complex agents with MAF:
- Middleware Composition: Chain function and chat middleware together to combine multiple handlers such as logging, authentication, and rate limiting
- Workflow Checkpointing: Leverage workflow events and serialization to save and resume long-running processes
- Dynamic Tool Selection: Combine RAG over tool descriptions with MAF tool registration to provide only relevant tools per query
- Multi-Agent Handoff: Handoff orchestration between specialized agents using workflow edges and conditional routing
Process FlowchartMermaid%%{init: {'look': 'handDrawn'}}%% flowchart TD MAF["Microsoft Agent\nFramework"] MAF --> Core["Core Concepts"] MAF --> WF["Workflows"] MAF --> Pattern["Advanced Patterns"] Core --> Agent["Agents\nCreate·Run·Tools·Threads"] Core --> MW["Middleware\nFunction·Chat"] Core --> Mem["Memory\nIn-memory·Persistent·Dynamic"] Core --> Obs["Observability\nOpenTelemetry"] WF --> Exec["Executor\nAI Agent or Custom Logic"] WF --> Edge["Edge\nDirect·Conditional·Fan-out·Fan-in"] WF --> Evt["Event\nStarted·Output·Error"] Pattern --> Seq["Sequential\nIterative Refinement"] Pattern --> Conc["Concurrent\nParallel Fan-out"] Pattern --> HO["Handoff\nDynamic Routing"] Pattern --> MWP["Middleware\nResult Override"] Pattern --> HITL["Human-in-the-Loop\nHuman Approval Integration"]