Cover Image for Stop Shipping AI Nobody Can Verify with Hamel Husain
Cover Image for Stop Shipping AI Nobody Can Verify with Hamel Husain
20 Going

Stop Shipping AI Nobody Can Verify with Hamel Husain

YouTube
Registration
Welcome! To join the event, please register below.
About Event

An AI data agent tells you net revenue was $4.21M. It does not show the metric definition, the query, the assumptions, or what it could not verify. The only way to trust the answer is to redo the analysis yourself.

That is a product problem.

Hamel Husain has spent the past three years focused on AI evals. When a product is hard to eval, users usually struggle to verify it too. Designing for verification comes before another scoring pipeline: people need to inspect the evidence, compare the output with something they trust, and review smaller units of work.

In this live conversation, Hugo Bowne-Anderson and Hamel will dig into a bigger product question: how should we design AI systems when verification is the bottleneck?

What We’ll Cover

• How to study the checks a domain expert performs before deciding what to evaluate

• Why an answer without provenance forces users to redo the agent’s work

• What data agents should expose: metric definitions, intermediate calculations, source queries, sanity checks, and unresolved inputs

• How vetted starting points and visible diffs reduce the amount of AI-generated work a person must review

• When an output generator should become a research assistant that links claims to evidence and surfaces contradictions

• Progressive disclosure: showing the evidence users need without burying them in every intermediate step

• Breaking AI output into smaller units that people can accept, edit, reject, or investigate

• How verification-friendly product design produces cheaper annotation and stronger eval signals

• Why verification is becoming the bottleneck in AI products

Register to join live or get the recording afterwards.

20 Going