AI Agents for Beginners - 10. AI Agents in Production: Observability & Evaluation

How to move AI agents from prototype to production — observability with traces and spans, key metrics, OpenTelemetry instrumentation, offline/online evaluation, cost management, and a real expense claim demo.
June 9, 2026

AI Agents for Beginners - 10. AI Agents in Production: Observability & Evaluation

This article summarizes Lesson 10 of Microsoft's AI Agents for Beginners course.
As AI agents transition from experimental prototypes to real-world applications, the ability to understand their behavior, monitor their performance, and systematically evaluate their outputs becomes critical.
The goal of this lesson is to provide the knowledge needed to transform a "black-box" agent into a transparent, manageable, and trustworthy system.

Traces and Spans

Traces and Spans are the core concepts for observing agent execution.
Observability tools such as Langfuse and Microsoft Foundry represent agent execution using this structure:
ConceptDescription
TraceThe entire agent task from start to finish (e.g., end-to-end processing of a user query)
SpanAn individual step within a trace (e.g., an LLM call, data retrieval, tool execution)
Trace & Span Hierarchy
Mermaid
flowchart TD T["Trace\nEntire Agent Task\n(User Query Processing)"] T --> S1["Span: LLM Call\nModel Inference"] T --> S2["Span: Tool Call\nget_flight_info()"] T --> S3["Span: Tool Call\nget_activity_suggestions()"] T --> S4["Span: LLM Call\nFinal Response Generation"] S1 -->|"latency, tokens"| M1["Metrics"] S2 -->|"latency, result"| M2["Metrics"] S3 -->|"latency, result"| M3["Metrics"] S4 -->|"latency, tokens"| M4["Metrics"]
Without observability, an AI agent feels like a "black box" where internal state and reasoning remain opaque.
With observability in place, the agent turns into a glass box, making it possible to build trust and verify that it is operating as intended.

Why Observability Matters in Production Environments

In production environments, observability is not a "nice-to-have" — it is a necessity:
ReasonDescription
Debugging & Root-Cause AnalysisPinpoint error root causes across complex multi-LLM calls, tool interactions, and conditional logic using traces
Latency & Cost ManagementPrecisely track latency and cost per API call to identify and optimize slow or expensive operations
Trust, Safety & ComplianceProvide audit trails for agent actions and decisions — detecting prompt injection, harmful content, and PII mishandling
Continuous Improvement LoopsEstablish feedback loops where production insights directly inform and improve offline experiments

Key Metrics to Track

Key metrics to track to monitor agent behavior:
MetricDescriptionExample Usage
LatencyAgent response speedAgent taking 20s → optimize with a faster model or parallel calls
CostsCost per agent executionIf 5 LLM calls yield marginal quality gains, reduce call count or use a cheaper model
Request ErrorsNumber of API errors and tool call failuresConfigure fallback to provider B when LLM provider A is down
User FeedbackExplicit ratings (👍/👎, ⭐1-5)Persistent negative feedback = signal that the agent is malfunctioning
Implicit User FeedbackRepeated questions, immediate rewrites, retry clicksAsking the same question repeatedly = signal that the agent did not deliver the expected response
AccuracyFrequency of correct and desirable outputsLabel traces with "succeeded"/"failed" → track success rates
Automated Evaluation MetricsLLM-based automated scoringRAGAS for RAG, LLM Guard for safety

Instrument your Agent

The goal of agent instrumentation is to make your code emit traces and metrics so they can be captured, processed, and visualized on an observability platform.

OpenTelemetry

OpenTelemetry (OTel) is establishing itself as the industry standard for LLM observability.
Microsoft Agent Framework integrates natively with OpenTelemetry:
otel_instrumentation.py
python

Manual Span Creation

While automatic instrumentation provides a baseline, you can manually create spans when more granular information is needed.
Add custom attributes such as user_id, session_id, and model_version to spans for debugging and analytics:
manual_span_langfuse.py
python

Building an Observable Agent

Implement simple observability by adding timing to a Travel Agent.
In production, integrate this with a tracing backend like OpenTelemetry:
setup.py
python
travel_tools.py
python
observable_agent.py
python
Observable Agent Flow
Mermaid
flowchart LR User["User Query\n'Plan a day trip in Paris'"] --> Timer["start_time = time.time()"] --> Agent["TravelAgent"] Agent --> F["get_flight_info('Paris')"] Agent --> A["get_activity_suggestions('Paris')"] F & A --> LLM["LLM\nSynthesize Response"] LLM --> Stop["elapsed = time.time() - start_time"] Stop --> Out["Response (elapsed s)\n+ Production: Send to OTel Backend"]

Agent Evaluation

Observability provides the metrics, while Evaluation is the process of analyzing that data to assess the agent's performance and determine how to improve it.
Because AI agents are non-deterministic and subject to change from updates and model drift, regular evaluation is essential.
Evaluation falls into two categories: Offline Evaluation and Online Evaluation.

Offline Evaluation

Dataset items in Langfuse
Evaluates the agent with a test dataset in a controlled environment.
Uses curated datasets with known expected outputs and runs during development (including CI/CD pipelines) to verify improvements or prevent regressions.
  • Advantages: Repeatable with ground truth, providing clear accuracy metrics
  • Key Challenges: Keeping test datasets comprehensive and continuously updated with real-world scenarios

Online Evaluation

Observability metrics overview
Evaluates the agent in a live, real-world environment.
Monitors performance during real user interactions and continuously analyzes outcomes.
  • Advantages: Captures phenomena unforeseen in lab environments — model drift, unexpected query patterns
  • Methods: Collecting implicit and explicit user feedback, shadow testing, A/B testing

Combining the two

Online and offline evaluations are not mutually exclusive — they are complementary:
Evaluation Loop
Mermaid
flowchart LR OE["Offline Evaluation\nEvaluate with Test Dataset"] --> Deploy["Deploy\nDeploy to Production"] --> OL["Online Monitoring\nTrack Real User Interactions"] --> Collect["Collect Failures\nGather New Failure Cases"] --> Add["Enrich Offline Dataset\nAdd Edge Cases & New Patterns"] --> Refine["Refine Agent\nImprove Prompts, Models, Logic"] Refine --> OE

Evaluation Patterns

A common production pattern is leveraging a second agent as an evaluator.
The evaluator agent scores the primary agent's responses against set criteria:
evaluator_agent.py
python
Evaluator Agent Pattern
Mermaid
flowchart TD User["User Query"] --> Primary["Primary Agent\n(TravelAgent)"] Primary --> Resp["Agent Response"] Resp --> Eval["Evaluator Agent\n(ResponseEvaluator)"] Eval --> Score["Scores\nCompleteness · Accuracy · Helpfulness · Overall"] Score --> Gate{Meets Quality Criteria?} Gate -->|"Pass"| Deliver["Deliver to User"] Gate -->|"Fail"| Flag["Flag → Refine / Retry"]

Common Issues

Common issues and solutions when deploying AI agents to production:
IssueSolution
Inconsistent Task ExecutionRefine prompts to clarify goals. Decompose tasks into subtasks and handle via multi-agent setups
Infinite LoopsDefine explicit termination conditions. Use reasoning-specialized large models for complex reasoning and planning tasks
Degraded Tool Call PerformanceTest and validate tool outputs outside the agent system. Refine parameters, prompts, and tool naming
Multi-Agent MisalignmentKeep each agent prompt clear and specific. Build hierarchical structures with "routing" or controller agents
With observability in place, these issues can be identified far more effectively. Traces and metrics help pinpoint the exact location of problems within the agent workflow.

Managing Costs

Cost management strategies for production AI agents:
StrategyDescription
Using Smaller ModelsSLMs perform well in specific agentic use cases and significantly reduce costs. Use SLMs for simple tasks (intent classification, parameter extraction) and large models only for complex reasoning
Using a Router ModelRoute requests to the optimal model based on complexity using an LLM/SLM or serverless function. Simple queries → small, fast models; complex reasoning → large models
Caching ResponsesCache common requests and tasks to serve responses before similar requests hit the agent system. Drastically reduces costs for FAQs and standard workflows
Cost Management: Router Model Strategy
Mermaid
flowchart TD Req["User Request"] --> Router["Router Model\n(LLM/SLM/serverless)"] Router --> Complexity{Classify Complexity} Complexity -->|"Simple\nIntent Classification & Parameter Extraction"| SLM["SLM\n(GPT-4o-mini, etc.)\nLow Cost · High Speed"] Complexity -->|"Complex\nComplex Reasoning & Planning"| LLM["Large Model\n(GPT-4o, etc.)\nHigh Performance"] Complexity -->|"Cached\nFrequently Asked Questions"| Cache["Cache\nInstant Response\nZero Cost"] SLM & LLM & Cache --> Response["Final Response"]

Lets see how this works in practice

As a hands-on example, we implement a receipt image OCR → expense claim email generation pipeline.
We connect an OCR Agent and an Email Agent using the multi-agent WorkflowBuilder.

Define Expense Models

expense_models.py
python

Defining Tools

expense_tools.py
python

Processing Expenses

expense_workflow.py
python
Expense Claim Workflow
Mermaid
flowchart LR Receipt["receipt.jpg\nReceipt Image"] --> OCR["OCRAgent\nload_receipt_image()\nImage → base64 Encoding"] OCR --> Parse["Parse Receipt Text\ndate|description|amount|category\nSemicolon-Separated Format"] Parse --> Email["EmailAgent\ngenerate_expense_email()\nExpenseFormatter.parse_expenses()"] Email --> Mail["Expense Claim Email\nDear Finance Team...\nTotal Amount: $xxx"]

Summary

Lesson 10 Summary
Mermaid
flowchart LR Root["AI Agents\nin Production"] Root --> Obs["Observability"] Root --> Eval["Evaluation"] Root --> Cost["Cost Management"] Obs --> O1["Traces & Spans"] Obs --> O2["Key Metrics\nLatency·Cost·Error·Feedback"] Obs --> O3["OpenTelemetry\n+ Manual Spans"] Eval --> E1["Offline Evaluation\nControlled Test Datasets"] Eval --> E2["Online Evaluation\nReal User Monitoring"] Eval --> E3["Evaluator Agent\nAutomated Quality Scoring"] Cost --> C1["Smaller Models (SLMs)"] Cost --> C2["Router Model"] Cost --> C3["Caching"]
  • Turn agents from "black boxes" into "glass boxes" using Traces & Spans, enabling effective debugging and optimization.
  • Track key metrics such as latency, cost, request errors, and user feedback to gain a comprehensive view of agent health.
  • Establish pre-deployment baselines with Offline Evaluation, and capture real-world drift and unexpected patterns with Online Evaluation.
  • Prevent regressions with the Evaluator Agent pattern by applying automated quality gating to primary agent responses.
  • Manage costs using router models, SLMs, and caching while reserving high-performance models for tasks that require them.
Jooojub
System S/W engineer
Explore Tags
Series
    Recent Post
    © 2026. jooojub. All right reserved.