

Eval Rubrics that Drive AI Product Strategy
The most under appreciated artifact every AI PM should invest in
Most teams treat evals as a testing chore: something you bolt on after the agent ships, using whatever generic quality score a tool gives you for free. They neglect to invest in custom eval rubrics - the most valuable AI artifact your team could create.
Your eval rubric is the definition of what "great" means for your product, and once you have it, it stops being a testing artifact and becomes infrastructure:
It's what coding agents can verify their work against to ship autonomously
It’s what tells you whether a cheaper open-source model can be swapped in, giving you more optionality for how you win with pricing your AI features
It’s what a reward model optimizes toward if you want to post train custom models
Teams without eval rubrics are flying blind on all three increasingly valuable fronts. This session is about how to build that rubric from scratch, for a real agentic product. You'll walk away with a framework you can apply immediately.
Who it's for: PMs, founders, and product leaders building or evaluating AI agents, no ML background required
Sandhya and Justin are the founders of Calibre, an applied AI research and consulting firm that works AI native startups as well as large enterprises looking to transform how they build products for the AI era.
Sandhya is an AI investor (Sequoia Capital, Khosla Ventures) and startup executive (Amplitude) with a background in software entrepreneurship and data science. She has led early rounds in high growth AI startups like AirOps and Vizcom.
Justin is a product leader and former founder with over 15 years of experience turning data and AI into scalable, customer-driven products. As Chief Product Officer at Amplitude, he helped pioneer the discipline of product analytics.