

Hosts: Junfan Zhu, Aurora Feng
discord.gg/WH7DrTHRXK
Robotics & World Models Reading Club 18: Is Latent All You Need for World Action Models? & Causal World Models. SF 07/18
Robotics & World Models Reading Club 18: Is Latent All You Need for World Action Models? From V-JEPA to DreamZero, FastWAM, and ImageWAM & Causal World Models For Real-World Intelligence — San Francisco 07/18
A high-signal reading group for AI researchers & builders pushing the frontiers of robotic world models, WAMs, and embodied intelligence. In our previous sessions, we brought together researchers and engineers from Boston Dynamics, Google DeepMind, NVIDIA, Stanford, UC Berkeley, Dyna, Physical Intelligence, Tesla, Generalist, Rhoda AI, and leading Bay Area robotics startups.
Hosted by Junfan Zhu & Aurora Feng.
Reading Club 18's Core Theme
Keynote 1: Is Latent All You Need for World Action Models? From V-JEPA to DreamZero, FastWAM, and ImageWAM
Keynote by: Guanming Wang & Bill (General Instinct, YC P26)
World Action Models (WAMs) have recently emerged as a promising paradigm for embodied AI, enabling robots to reason about future observations and actions jointly. While early approaches such as DreamZero rely on generative video prediction to learn rich world representations, more recent methods like FastWAM and ImageWAM suggest that explicitly generating future videos may not be necessary. Instead, predicting and reasoning over latent world representations could be sufficient for effective action generation.
In this talk, we revisit this question through the lens of representation learning. Starting from V-JEPA, we discuss the philosophy of predictive latent representations, followed by DreamZero, which unifies future video and action prediction, and finally recent efficient WAMs including FastWAM and ImageWAM, which increasingly shift computation from pixel generation to latent reasoning. We will compare their design choices, discuss the trade-offs between latent prediction and video generation, and explore what information a latent world representation must contain to support robust robotic decision making.
Rather than presenting individual papers in isolation, this talk aims to provide a unified perspective on the evolution of World Action Models and examine an open research question for the community:
Do robots really need to generate future videos, or is learning the right latent representation enough?
Keynote 2: Towards Causal World Models for Physical AI
Aether AI has raised $20M to build causal world models that understand mechanisms.
Keynote by Biwei Huang (Professor at UCSD) & Feng Fan.
World models have emerged as a key foundation for physical AI by enabling agents to predict future observations, simulate outcomes, and plan before acting. However, current world models often struggle with causal reasoning, physical consistency, long-horizon decision making, and action understanding. In this talk, I will argue that the next generation of world models should move beyond predictive dynamics toward causal world models that capture the underlying mechanisms governing the physical world. I will present recent advances in three directions: learning causal representations of hidden state factors, discovering latent actions and skills from interaction, and building self-improving world models through causal feedback. Together, these developments point toward a new foundation for more robust, generalizable, and reliable physical AI.
Location
Studio 45 - informal spaces SF is a physical space for entrepreneurs and teams building hardware in SF. Located in Mission/Bernal Area, the studio gathers the community, resources, space, and tools needed to build a business making physical products.
Date & Time
Saturday, July 18, 2026 | 2:00 PM – 5:00 PM
Join Discord Community
https://discord.gg/WH7DrTHRXK
Follow Saturday Robotics on X
https://x.com/saturdayrobotic
Agenda
2:00 PM – 2:30 PM Door Opens & Social
Food 😋, beverages🧋 and UNLIMITED strawberries 🍓 (our official reading club fruits ☺️😄).
2:30 PM – 4:30 PM Keynote 1 by Guanming Wang & Bill (General Instinct, YC P26)
Keynote 2 by Biwei Huang, Professor UCSD (https://biweihuang.com/)
Online access via Zoom: TBD
YouTube Recording: TBD (We are looking for recording volunteers)
4:30 PM – 5:00 PM Q&A, open-floor roundtable (10–20 min per topic) on spotlight papers or any paper you’d like to highlight. Feel free to share why the paper matters and its technical details.
Future events
#reading-club-20-0725: Agentic Robotics Models (ENPIRE and Cap-X), Mountain View 07/25
Session 20 Luma: https://luma.com/5ltk12w5
#SIGGRAPH-reading-club-19-0722: SIGGRAPH x Saturday Robotics — World Models for Robotics: Bridging Graphics, Simulation & Physical Intelligence | Reading Club 19, LA 07/22
Session 19 Luma: https://luma.com/yh2212ac
Past events
#reading-club-17-0711: Soft Tactile-Centric Multimodal Intelligence Toward Safe and Dexterous Manipulation. SF 07/11
Session 17 Luma: https://luma.com/e53zawq2
#reading-club-16-0704: The Embodied AI Hardware Stack — Supply Chain, Sensors, and the Data Flywheel — SF 07/04
Session 16 Luma: https://luma.com/cgzyfpeb
#reading-club-15-0627: Scaling Touch: Flexible Tactile Skin for Dexterous Manipulation
Binghao Huang, Columbia, Amazon FAR.
Session 15 Luma: https://luma.com/7lts5ppf
Reading Club 15 Review: https://x.com/junfanzhu98/status/2071129728984273105?s=20
Reading Club 15 Recap: https://x.com/junfanzhu98/status/2071132114226139489?s=20
#deep-tech-week-14-0625: Deep Tech Week Research Night
SPEAR: A Simulator for Photorealistic Embodied AI Research. (ECCV 2026 accepted) by Mike Roberts, Senior Research Scientist, Adobe Research.
Bogdan Cristei, Venture Partner at SHACK15 Ventures.
Simone Totaro, CTO at Saturn Dynamics.
Shumo Chu, CEO at General Intelligence Labs.
Margaret Zhang, CEO at ThirdBrain Labs.
Session 14 Luma: https://luma.com/l1g9c2l1
Deep Tech Week 14 Review: https://x.com/junfanzhu98/status/2070417784560066592?s=20
Deep Tech Week 14 Recap: https://x.com/junfanzhu98/status/2070420270964514923?s=20
#reading-club-13-0620: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620
Session 13 Luma: https://luma.com/6vkhxnum
Reading Club 13 Review 1: HumanEgo: Zero-Shot Robot Learning, Human Egocentric Video, Leo Wang, Amazon FAR & UMaryland https://x.com/junfanzhu98/status/2068603103138713824?s=20
Reading Club 13 Recap 1: https://x.com/junfanzhu98/status/2068605511549743127?s=20
Reading Club 13 Review 2: Causal World Models: Biwei Huang, Aether AI & UCSD
Reading Club 13 Recap 2: https://x.com/junfanzhu98/status/2068746229643936177?s=20
#reading-club-12-0613: Origami Robotics (YC W26) on Dexterity
Session 12 Luma: https://luma.com/5w7c1t2a
Reading Club 12 Review: https://x.com/junfanzhu98/status/2066275974988337178?s=20
Reading Club 12 Recap: https://x.com/junfanzhu98/status/2066278639554245097?s=20
#cvpr-denver-11-0606: 🤖🥘 Saturday Robotics x Manycore Tech x Neural Motion | CVPR 2026 Denver Research Night | Robotics & World Models Reading Club 11
Junfan Zhu & Aurora Feng, Founders of Saturday Robotics
Anthony Zhao, Head of North America at Manycore Tech SpacialVerse
Aurora Feng, Founder at Neural Motion. NM-GenET.
Max Zhaoshuo Li, Robotics and World Model Tech Lead at NVIDIA Cosmos. Cosmos 3.
Xiaofan Li, World Model Tech Lead at X Square Robot. WALL-WM.
Zesen Zhao, University of Michigan. Test-Time Scaling for World Action Models via Zero-Shot Geometric Verification.
Pengyi Liao. VGGT-Ω: From 3D Reconstruction to Scalable Spatial Representation.
Jie Wang, University of Pennsylvania, GRASP Lab. Toward a Robotics MMLU: Lessons from Sim & Real Evaluations of Generalist Policies.
Gordon Qian, Senior AI Researcher at Snap. Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning.
YouTube livestream: https://www.youtube.com/live/P_3gSC-5cYM?si=530zqf2NscCq643O
CVPR Denver Research Night Luma: https://luma.com/zamm9g2g
CVPR Denver Research Night Lightning Talks Review: https://x.com/junfanzhu98/status/2065102150418788581?s=20
CVPR Denver Research Night Lightning Talks Recap: https://x.com/junfanzhu98/status/2065104616497524843?s=20
CVPR Hot Takes Review: https://x.com/junfanzhu98/status/2065234892653547904?s=20
CVPR Hot Takes Recap: https://x.com/junfanzhu98/status/2065236892988416166?s=20
#reading-club-10-0530: Bringing Robots to Life — Learning Humanoid Instincts from the Body Up | San Francisco 0530
Haochen Shi (Stanford, co-advised by Karen Liu & Shuran Song)
Session 10 Luma: https://luma.com/czz76qe1
Reading Club 10 Recap: https://x.com/junfanzhu98/status/2061145697693683878?s=20
#private-dinner-01-0529: Robo Plov x Saturday Robotics
Private Dinner 01 Luma: https://luma.com/3rzqwond
Private Dinner 01 Review: https://www.linkedin.com/posts/junfan-zhu_saturday-robotics-was-excited-to-host-its-activity-7466384182190182400-Zgnv?utm_source=share&utm_medium=member_desktop&rcm=ACoAABxP-p0BpUNGDf347aKh_1uJAPzG4er0As8
#reading-club-09-0523: CVPR Warm-up & Founders Spotlight — DeltaWorld + VisuoTactile Dexterous Hands
Tommie Kerssies (Amazon Frontier AI & Robotics)
Arjun Subramaniam (Factory Intelligence)
Session 09 Luma: https://luma.com/wooiz0bf
Reading Club 09 Review: Part 1: Robotics & World Model Reading Club 9.1: CVPR Warm-up— A Frame is Worth 1 Token: DeltaToken. https://x.com/junfanzhu98/status/2058449627184267621?s=20
Reading Club 09 Review: Part 2: Robotics & World Model Reading Club 9.2: Tactile Sensor & Reliable Manipulation in Production. https://x.com/junfanzhu98/status/2058461947637694948?s=20
#reading-club-08-0516: Embodied Human Data as the “Internet of Motion and Behavior”
Ryan Punamiya (NVIDIA Gear, Georgia Tech)
Session 08 Luma: https://luma.com/qoxioge7
Reading Club 08 Review: https://x.com/junfanzhu98/status/2055915875493204439?s=20
#reading-club-07-0509: Learning to Dream: World Models, Imagination, Path to Foundation Models for Control
Ahmet Şemi ASARKAYA (Agility Robotics)
Session 07 Luma: https://luma.com/srhe0vuo
Reading Club 07 Review: https://x.com/junfanzhu98/status/2053387034241454397?s=20
#reading-club-06-0502: Evolution of Video World Models for Robotics
Session 06 Luma: https://luma.com/sdrd4zwr
Reading Club 06 Review: https://x.com/junfanzhu98/status/2050834699275383008?s=20
#reading-club-05-0425: World Models for Physical Intelligence: From Predictive Brains to Embodied Robots
Daniel Dugas & Sergio Arnaud (Meta FAIR)
Session 05 Luma: https://luma.com/p7zvpyvg
Reading Club 05 Review: https://x.com/junfanzhu98/status/2048315020946317710?s=20
YouTube Recording: https://youtu.be/RVy6oQXNDgc?si=u2VLtCBjfdMvXaf-
#reading-club-04-0418: Abstractions of the Physical World for Decision-Making
Siming He (UC Berkeley)
Session 04 Luma: https://luma.com/atv7bm3i
Reading Club 04 Review: https://x.com/junfanzhu98/status/2045770010979905862
YouTube Recording: https://www.youtube.com/@saturdayrobotic
#reading-club-03-0411: Robotic Policy Adaptation
Haoyi Niu (UC Berkeley)
Session 03 Luma: https://luma.com/561xgirg
Reading Club 03 Review: https://x.com/junfanzhu98/status/2043243484568768519?s=20
YouTube Recording: https://www.youtube.com/@saturdayrobotic
#reading-club-02-0404: JEPA Zoo
Julian Saks (https://x.com/JulianSaks)
Session 02 Luma: https://luma.com/g3qrrti0
Reading Club 02 Review (liked by Yann LeCun on X): https://x.com/junfanzhu98/status/2040716119259164673?s=20
#reading-club-01-0328
Session 01 Luma: https://luma.com/8s4w1wu6
Reading Club 01 Review (liked by Yann LeCun on X): https://x.com/junfanzhu98/status/2038153945219305812
Logistics
Spots are limited. Please arrive by 2:00 PM for check-in. Keynote will begin promptly at 2:30 PM.
We currently do not have volunteers available to assist with late check-ins. Given the high volume of inquiries and 100+ attendees (both online and onsite), we kindly ask that you arrive on time to ensure smooth entry.
Hosts: Junfan Zhu, Aurora Feng
discord.gg/WH7DrTHRXK