

Scaling voice agents to Millions calls/week: Probabilistic models, deterministic outcomes
“You have some elbow pain, so you call ACME Clinic. An AI answers, and you’re impressed: it understands you perfectly, finds an available slot, and books you with a new doctor, Dr. Smith. You hang up thinking, Wow, that was easy.
The following week, you show up at the clinic—only to discover that Dr. Smith is actually a podiatrist. He specializes in feet. Not elbows. Now you have to reschedule for next week. You wasted a trip, the clinic wasted a slot, and you’re probably never trusting an AI to book a medical appointment again”.
LLMs are great at understanding people and terrible at guaranteeing anything. Production voice agents need both, combined with real-time latency. This talk explores how we combine LLMs with structure to turn nondeterministic models into reliable production systems.
We’ll explore which mechanisms can be used in high-stakes domains like medical and emergency scenarios, and why orchestrating fast models, reasoning models, and parallel perception systems is a non-trivial problem that demands new ideas.
Agenda
Prosper Pitch + Business/Product Context — 5 minutes
Tech Talk — 20 minutes
Scheduling complexity
Many complex custom rules
Hard constraints
Key features
Gates
Custom code
Parallel LLMs
Jinja
Structured Outputs
Open Q&A — 10 minutes
Networking & Beers