

Presented by
Braintrust
Braintrust is the active observability platform for agents.
Hosted By
Measure what matters: Intro to AI evals for common use cases
Registration
Registration Closed
This event is not currently taking registrations. You may contact the host or subscribe to receive updates.
About Event
This session will break down the basics through practical examples that are suitable for both AI engineers and PMs.
We'll work through how to evaluate three common use cases:
Customer support agent: "Is my AI support agent ready for customers?"
Content/code generation: "How do I know if my AI's output is actually good?"
Prompt optimization and model testing: "Which version of my AI setup works better?"
We'll also cover how to write good scoring functions and manage datasets. No prior evaluation experience is required. Framework and model-agnostic approaches that work with any AI application.
Presented by
Braintrust
Braintrust is the active observability platform for agents.
Hosted By