Cover Image for BRISA Reading Group 02 - Do WAMs Generalize Better than VLAs?
Cover Image for BRISA Reading Group 02 - Do WAMs Generalize Better than VLAs?
Avatar for BRISA
Presented by
BRISA
Berlin Robotics & Intelligent Systems Association
Building the future of Physical AI in Berlin
13 Going

BRISA Reading Group 02 - Do WAMs Generalize Better than VLAs?

Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

Before our official launch this fall, we’re continuing our pre-launch series with another reading group: sitting down with a paper, understanding it properly, and asking what it would take to build it ourselves.

Our second paper is Do World Action Models Generalize Better than VLAs? A Robustness Study (Zhang et al., 2026), a comparative study of state-of-the-art Vision-Language-Action models and the emerging class of World Action Models.

This time, we’re looking at a question that is becoming increasingly important in robot learning: does explicitly predicting how the world evolves make robot policies more robust?

VLAs map visual observations and language instructions directly to robot actions. World Action Models instead build on models trained to predict how the visual world evolves over time, using those learned dynamics to support action generation.

The paper asks whether that difference actually leads to better generalization.

The authors compare WAMs and VLAs across LIBERO-Plus and RoboTwin 2.0-Plus, testing them under changes in camera viewpoint, robot state, language, lighting, background, noise and object layout.

The results are interesting, but not one-sided. WAMs are particularly robust under several visual perturbations, while strong VLAs remain highly competitive in parts of the evaluation.

That makes the paper less about declaring a winner and more about a broader question for Physical AI:

Should we keep scaling models that map perception directly to action, or should a robot first learn to imagine what happens next?

The paper also gives us a concrete robustness setup to discuss, rather than just architecture diagrams and benchmark scores.

No preparation needed. Skim the abstract if you have ten minutes. Everyone is welcome — whatever you study, wherever you study and however much robotics you have done so far.

Come with questions, disagreements, or your own take on where robot foundation models should go next.

🎟️ There are only a limited number of spots, so we can keep the session personal and interactive. Please register early and only attend if you've received a ticket.

📍 Hosted at the Merantix AI Campus. Doors open 18:30, we start at 18:45.

👋 New here? Join the community group — everything else gets announced there first.

Location
Merantix AI Campus
Max-Urich-Straße 3, 13355 Berlin, Germany
Avatar for BRISA
Presented by
BRISA
Berlin Robotics & Intelligent Systems Association
Building the future of Physical AI in Berlin
13 Going