

AI Evals and Observability Strategies for Web Developers
Building an AI agent for your application is only the beginning. The harder challenge is knowing whether it is actually working well once it starts planning, calling tools, making decisions, and operating across multiple steps.
This workshop shows you how to design a practical evaluation and observability strategy for agentic systems. You’ll learn what to measure, how to evaluate agents during development, how to assess multi-step trajectories, and how to use production telemetry and online evals to detect failures and improve reliability.
The goal is to help you move from “the agent seems to work” to a repeatable, evidence-based approach for measuring agent quality.
What you’ll learn
By the end of the workshop, you’ll be able to:
Identify what should be evaluated in an agentic system
Design meaningful test cases and evaluation criteria
Run offline evals before deploying agent changes
Evaluate tool use, reasoning paths, and multi-step execution
Use traces and production signals to diagnose agent failures
Introduce online evals to continuously monitor quality
Build a repeatable evaluation strategy for your own agents
Meet Your Instructor
Supreet Kaur
Sr. Gen AI Solutions Architect | AWS
Supreet is a Senior GenAI Solutions Architect at AWS, helping startups take AI ideas from proof of concept to production. Her career spans data science, financial services, and cloud AI architecture, with roles at ZS Associates, Morgan Stanley, Microsoft, and AWS. She is the author of The AI Optimization Playbook, a speaker at 40+ events, a published thought leader, and co-inventor of a patented AI-powered personalization testing strategy.
What you’ll leave with
You’ll receive:
A practical framework for evaluating AI agents
Guidance on choosing the right agent quality metrics
Approaches for offline evaluation during development
Techniques for assessing agent trajectories and intermediate decisions
A clearer understanding of observability for production agents
Strategies for combining offline and online evaluation
Best practices for building an eval strategy that evolves with your agent
Full workshop recording
Certificate of completion
More importantly, you’ll leave with the vocabulary and intuition to discuss AI systems more precisely, investigate failures more effectively and make better decisions when building AI-powered products.
Who should attend?
This workshop is ideal for:
Web and Software Developers building agentic applications
AI Engineers
Platform and Backend Engineers
Technical Leads and Architects
Teams moving AI agents from prototype to production
It is particularly useful for anyone who already has an agent or agentic workflow and wants a more systematic way to measure, debug, and improve it.