

Live workshop: Bringing evals to production
Registration
Past Event
About Event
In this session, you'll instrument an app with Braintrust logging and configure online scoring so your scorers run automatically on real traffic. Then you'll pull interesting production rows (edge cases you didn't anticipate, inputs that scored poorly) directly into your eval dataset. This is the cycle where production logs surface issues, issues become test cases, and test cases drive system iteration.
You should walk away with a concrete understanding of how evals connect to production, and how to build a feedback loop that makes your AI feature measurably better over time.