

Evaluate Any AI Agent: Pre-Production to Production (platform credits will be offered at the event)
Hands-on workshop for AI engineers. Bring any agent (RAG, voice, MCP, coding copilot, browser, anything) and leave with evals that catch failures before users do - No additional API keys required, just bring your agent.
Your agent works in your demo. In production, it hallucinates, mis-calls tools, breaks workflows, and fails silently. Only at scale. Only on edge cases. Only in front of real users.
This workshop is built to fix that.
In 2 hours, we walk through evaluating any AI agent across its full lifecycle. You bring an agent. We help you wire up real evals: before launch, in production, and the loop that connects them.
Pre-Production Evaluation
Before users ever touch your agent:
• Generate realistic and adversarial test conversations
• Simulate edge cases at scale
• Measure hallucination rates and task success
• Evaluate tool-calling correctness
• Stress-test workflows before they ship
• Build and deploy custom evaluators specifically for your agent
Production Evaluation
Once your agent is live:
• Monitor real-world failures
• Track reliability and regressions over time
• Build feedback loops for continuous improvement
• Detect failure patterns before they become incidents
Closing the Loop: Optimization from Eval Data
Evals are not just measurement. They feed back into product improvement:
• Pull insights from production traces
• Identify what to optimize, prioritize, and fix
• Walk out with an eval pipeline tailored to your stack
Bring Your Own Agent
Framework-agnostic. The workshop is built around evaluation principles that work across stacks.
Examples: OpenAI Agents SDK, LangChain, CrewAI, MCP-based systems, RAG pipelines, browser agents, voice agents, internal copilots, early prototypes, anything.
Who This Is For
• AI engineers shipping agents
• Founding engineers
• Backend developers working with LLMs
• AI product teams
• Developers debugging weird LLM failures in production
Hands-on technical session for builders, not a lecture.
What to Bring
• Laptop (compulsory)
• An agent or prototype you want to evaluate
• A willingness to break things and fix them
Agenda
1. Why AI agents fail in production
2. Pre-production simulation workflows
3. Building structured eval pipelines
4. Catching hallucinations and tool-call errors
5. Production monitoring and observability
6. Closing the loop: optimization from eval data
7. Live demos, Q&A, networking + refreshments
The Tools
We use Future AGI's open-source stack:
• Simulation
• Evaluation
• Optimization
• Observability
• Guardrails
• Production monitoring
You leave with workflows you can apply Monday morning.