Cover Image for BRISA Reading Group 04 - LeWorldModel
Cover Image for BRISA Reading Group 04 - LeWorldModel
Avatar for BRISA
Presented by
BRISA
Berlin Robotics & Intelligent Systems Association
Building the future of Physical AI in Berlin

BRISA Reading Group 04 - LeWorldModel

Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

Before our official launch this fall, we’re continuing our pre-launch series with another reading group: sitting down with a paper, understanding it properly, and asking what it would take to build it ourselves.

Our fourth paper is LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels (Maes et al., 2026), which asks:

How simple can a world model become while still being useful for planning?

World models learn to predict how an environment might change under different actions, making it possible to compare possible futures before acting.

LeWorldModel, or LeWM, takes a relatively simple approach. Instead of predicting future images pixel by pixel, it learns a compact representation of the scene and predicts how that representation changes when different actions are taken.

One of the main challenges with this kind of model is preventing those learned representations from collapsing during training. LeWM uses SIGReg to stabilize this process, allowing the whole model to be trained directly from pixels and actions without relying on a pretrained visual encoder.

The resulting model has around 15 million parameters, can be trained on a single GPU in a few hours, and uses its learned dynamics to plan toward a goal.

The authors report planning up to 48× faster than DINO-WM, while remaining competitive across several of the evaluated control tasks.

What makes the paper especially interesting is how little machinery it appears to need. It raises the possibility that relatively small world models may already be enough to learn useful dynamics and support effective planning.

At the reading group, we’ll look at why this relatively simple approach works, what the model actually learns about its environment, and where that simplicity may start to become a limitation.

The broader question is:

How far can a compact latent world model like this scale toward real robotic manipulation?

No preparation needed. Skim the abstract if you have ten minutes. Everyone is welcome — whatever you study, wherever you study and however much robotics you have done so far.

Come with questions, disagreements, or your own take on what a world model for robotics should actually learn.

🎟️ There are only a limited number of spots, so we can keep the session personal and interactive. Please register early and only attend if you've received a ticket.

📍 Hosted at the Merantix AI Campus. Doors open 18:30, we start at 18:45.

👋 New here? Join the community group — everything else gets announced there first.

Location
Merantix AI Campus
Max-Urich-Straße 3, 13355 Berlin, Germany
Avatar for BRISA
Presented by
BRISA
Berlin Robotics & Intelligent Systems Association
Building the future of Physical AI in Berlin