

Building the future of Physical AI in Berlin
BRISA Reading Group 03 - TOPReward
Before our official launch this fall, we’re continuing our pre-launch series with another reading group: sitting down with a paper, understanding it properly, and asking what it would take to build it ourselves.
Our third paper is TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics (Chen et al., 2026), which asks a surprisingly simple question:
Can a pretrained Vision-Language Model already tell us how well a robot is doing — without training a separate reward model?
In robot learning, knowing whether a task was completed is relatively easy. Knowing how much progress the robot is making along the way is much harder.
TOPReward approaches this without manually designing rewards or training another model.
Instead of asking a VLM to explicitly output a progress score, it asks the model whether a task has been completed and looks directly at the probabilities it assigns to the possible answer tokens.
Those token probabilities become a zero-shot reward signal that can be used to measure task progress, detect success, and reweight robot demonstrations.
What makes the idea especially interesting is that the reward model itself was never explicitly trained to be a reward model.
The paper therefore raises a broader question:
How much useful supervision is already hidden inside foundation models if we know where to look for it?
TOPReward also entered our exploration through MolmoAct2, where it is used as a reward signal inside a larger robot-learning pipeline.
During the session, we’ll focus on three things: how TOPReward actually works, how it is applied to a relatively simple visuomotor policy like ManiFlow — which is not a VLA — and what happens to policy performance when demonstrations are reweighted using this signal.
That gives us an interesting question beyond TOPReward itself:
Can knowledge hidden inside a VLM improve robot policies that do not use a VLM at inference time at all?
And, as always, we’ll look beyond the headline results: how convincing is the reward signal, what does it actually capture, where might it fail, and how difficult would it be to reproduce the idea ourselves?
No preparation needed. Skim the abstract if you have ten minutes. Everyone is welcome — whatever you study, wherever you study and however much robotics you have done so far.
Come with questions, disagreements, or your own take on how foundation models should be used in robot learning.
🎟️ There are only a limited number of spots, so we can keep the session personal and interactive. Please register early and only attend if you've received a ticket.
📍 Hosted at the Merantix AI Campus. Doors open 18:30, we start at 18:45.
👋 New here? Join the community group — everything else gets announced there first.
Building the future of Physical AI in Berlin