

Evals for Physical AI | Robohouse x Positronic
Today it is hard to tell how good these models are outside a demo. We believe measurement is one of the most important needs in this field.
Most real-robot evals still come down to binary success rate, a fixed timeout, and 25 rollouts or fewer, usually with no confidence intervals. That's not enough to distinguish close models.
In this session, we dig into PhAIL (Physical AI Leaderboard), an open real-robot benchmark on a Franka FR3 that replaces the success-rate coin flip with time-to-success distributions, human-relative scoring, and proper significance tests. On four public VLAs, it separates pairs that binary metrics can't. One pair still can't be separated. And the best model is roughly 7× slower than a human doing the same job.
Agenda
5:30 PM: Lecture and open discussion with the founder of Positronic
6:30 PM: BBQ
We'll get into:
Why success rate at a fixed timeout hides most of what matters
Time-to-success CDFs and Human-Relative Throughput
What N≤30 rollouts can and can't resolve
What this means if you're training, buying, or deploying robots
Part of the Random Intimate Lectures series at Robo House, for robotics engineers, researchers, and anyone deploying real robots. Small room, real numbers. Then we eat.
Paper: arxiv.org/abs/2605.29710 | Leaderboard: phail.ai