How do you know your agent works when it scales beyond 1K sessions?
Agents at Scale — a meetup series on running AI agents in production, from the team that ran one through the 2026 World Cup measuring 56 million devices online at once.
This session is about measuring what users did rather than scoring what the agent said, drawn from what we learned running an analytics agent under live World Cup load. The agent was pressure-tested by 20 major streaming providers using it to analyze live streaming data and ask questions about their viewer’s quality of experience in real time.
What we'll cover
The blind spots we had using traditional evals
The implicit signals we measured instead: clarification ratio, repetition rate, scope-narrowing, verification-seeking, post-conversation correction and escalation
How we use population-level contrastive analysis to stop asking which trace failed, and starting seeing under what conditions it's failing
Three metrics you can instrument this week from data you already have
Who this is for
Engineers and product owners with an agent already in production, especially if your traffic is spiky, your users ask things you didn't anticipate, or your agent passes all of your evals but you’re still losing customers after they interact with your agent.
Not an intro session. We'll assume you've shipped a consumer-facing agent (or about to ship one).
Details
Sept 24, 2036. 6:00 PM, Foster City (Full Address on registration)
Doors 6 pm, talk at 6:30 pm, then open Q&A — stick around, the hallway conversation is half the point. Food and drinks provided. Limited to spots.
Speakers:
Haijie Wu - Sr. Director, Product Management, Conviva
Junhwi Kim - AI Lead, Engineering. Conviva
Hosted by Conviva. We've spent twenty years measuring whether video actually worked for the person watching it. We're now doing the same for AI agents, and we'd rather compare notes with people building them than talk at them.
