

Live workshop: Multi-turn scoring
Registration
Past Event
About Event
In this workshop, you'll learn the difference between single-turn and multi-turn scoring. Then you'll build a trace-level scorer that evaluates full conversations as a unit, and set up online scoring so both levels of signal run automatically on every new log. You'll instrument a multi-turn chat app with production logging and see how the two scores complement each other to help you understand your AI system as a whole.
You should walk away understanding when single-turn evals are sufficient, when they're not, and how to set up scoring that catches the failures that only show up across multiple turns.