Cover Image for Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29
Cover Image for Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29
Avatar for Saturday Robotics
Presented by
Saturday Robotics
🤖 Saturday Reading Club on Robotics & World Models for AI Researchers in SF
Hosts: Junfan Zhu, Aurora Feng
discord.gg/WH7DrTHRXK
84 Went

Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap+Video-Tactile-Action Model. SF 8/29

Register to See Address
San Francisco, CA
Registration
Past Event
Please click on the button below to join the waitlist. You will be notified if additional spots become available.
About Event

Robotics & World Models Reading Club 26: Video Generation to Robot Manipulation: Bridging Embodiment Gap + Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs. SF 8/29

A high-signal reading group for AI researchers & builders pushing the frontiers of robotic world models, WAMs, and embodied intelligence. In our previous sessions, we brought together researchers and engineers from Boston Dynamics, Google DeepMind, NVIDIA, Stanford, UC Berkeley, Physical Intelligence, Tesla, Generalist, Rhoda AI, and leading Bay Area robotics startups.

Hosted by Junfan Zhu & Aurora Feng.

​​​Reading Club 26's Core Theme

Keynote 1: From Video Generation to Robot Manipulation: Bridging the Embodiment Gap

Keynote: Karthik Dharmarajan (UC Berkeley)

Recent advances in video generation models enable zero-shot synthesis of physically plausible object interactions, offering a compelling prior for open-world robot manipulation. However, these models operate in pixel space and work better with human embodiments, making it challenging to translate the generated videos into actionable robot behavior. In this talk, I will present an approach that uses 3D object flow as the interface between video generation and robot control. By extracting object-centric flow directly from generated videos or 4D (3D + time) representations, we convert visual predictions into a reward signal for trajectory optimization and reinforcement learning. This formulation decouples what should happen (object state changes) from how a robot embodiment should achieve those changes, allowing frameworks like Dream2Flow and CHORD to overcome the human–robot embodiment gap. Together, these methods enable zero-shot guidance from pre-trained video models to manipulate a wide range of objects including rigid, articulated, deformable, and granular.


​​​​​​​​​​​​​Pre-Readings

Dream2Flow: https://arxiv.org/pdf/2512.24766

CHORD: https://arxiv.org/pdf/2601.04194


Keynote 2: VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Keynote: Haoran Yuan (UC Berkeley, Applied Intuition)

Recent vision-language-action models commonly incorporate tactile observations through direct feature injection, enabling policies to respond to contact feedback only after physical interaction has occurred. This presentation introduces VTAM, a video-tactile-action framework that instead models the future evolution of both visual and tactile observations. VTAM adopts a continual-learning formulation in which a pretrained visual world model is extended to tactile dynamics through fine-tuning, without requiring a separately pretrained tactile encoder or tactile representation model. By leveraging the shared visual structure of RGB observations and vision-based tactile signals, the model learns a unified visuo-tactile latent dynamics space while preserving the knowledge acquired during visual pretraining. The predicted visuo-tactile dynamics are subsequently integrated into the action model to support contact-aware action generation. Compared with direct tactile injection, which primarily provides reactive feedback from the current observation, tactile world modeling enables the policy to anticipate future contact states, force transitions, and potential slip before action execution. This predictive capability provides a foundation for proactive tactile planning and more robust manipulation in contact-rich environments.


​​​​​​​​​​​​​Pre-Readings

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

https://arxiv.org/pdf/2603.23481


Location

San Francisco

​​​​​​​​​​​​​Date & Time

Saturday, August 29, 2026 | 2:00 PM – 5:00 PM

​​​​​​​​​​​​​​Join Discord Community

Join Discord Server

​​​​​​​​​​​​​Follow Saturday Robotics on X

/saturdayrobotic


​​​​​​​​​​​​​Agenda

2:00 PM – 2:30 PM Door Opens & Social

  • Food 😋, beverages🧋 and UNLIMITED strawberries 🍓 (our official reading club fruits ☺️😄).

2:30 PM – 4:30 PM Keynote by Karthik Dharmarajan (UC Berkeley), Haoran Yuan (UC Berkeley, Aplied Intuition)

YouTube Recording: TBD (We are looking for recording volunteers)

4:30 PM – 5:00 PM Q&A, ​open-floor roundtable (10–20 min per topic) on spotlight papers or any paper you’d like to highlight. Feel free to share why the paper matters and its technical details.


Future events

#iros-reading-club-31-0928: 🍾 IROS 2026 x Saturday Robotics — Robotics Research Night | Reading Club 31. Pittsburgh 9/28


​​​​​​​​​​​​​​​​​​​Past events

#reading-club-26-0905: Rethinking Robot Development: Co-Designing Morphology, Sensing, and Learning

Chenyang Ma (Applied Intuition, UNC)

#reading-club-25-0829: Sunnyvale 8/29

#reading-club-24-0822: Contact-Rich Robot Learning from Human Videos and Tactile. SF 8/22

Kelin Yu (Maryland, Amazon FAR)

#reading-club-23-0815: Engineering Robotic Simulators for Evaluation and Beyond. SF 8/15

Kaifeng Zhang (Columbia, World Labs)

#reading-club-22-0808: ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow. SF 8/8

Haoyi Niu (UC Berkeley)

#reading-club-21-0801: Vision-Language-Kinematics Supervision for Perception-Based Humanoid Loco-Manipulation — SF 8/1

Yen-Jen Wang (UC Berkeley, Amazon FAR)

#reading-club-20-0725: Agentic Robotics Models (ENPIRE and Cap-X), Mountain View 07/25

Haoru Xue (UC Berkeley)

Kris Hauser (Samsung Research America, Robot Intelligence Lab)

#SIGGRAPH-reading-club-19-0722: SIGGRAPH x Saturday Robotics — World Models for Robotics: Bridging Graphics, Simulation & Physical Intelligence | Reading Club 19, LA 07/22

#reading-club-18-0718: Causal World Models For Real-World Intelligence. SF 07/18

Guanming Wang & Bill (General Instinct, YC P26)

Feng Fan (UCSD, Aether AI)

#reading-club-17-0711: Soft Tactile-Centric Multimodal Intelligence Toward Safe and Dexterous Manipulation. SF 07/11

Quan Luu, Purdue.

#private-lunch-icml-0709: Saturday Robotics x ICML Private Lunch (Seoul)

#reading-club-16-0704: The Embodied AI Hardware Stack — Supply Chain, Sensors, and the Data Flywheel — SF 07/04

Jerry Huang, Robotics Center of Silicon Valley.

#reading-club-15-0627: Scaling Touch: Flexible Tactile Skin for Dexterous Manipulation

Binghao Huang, Columbia, Amazon FAR.

#deep-tech-week-14-0625: Deep Tech Week Research Night

SPEAR: A Simulator for Photorealistic Embodied AI Research. (ECCV 2026 accepted) by Mike Roberts, Senior Research Scientist, Adobe Research.

Bogdan Cristei, Venture Partner at SHACK15 Ventures.

Simone Totaro, CTO at Saturn Dynamics.

Shumo Chu, CEO at General Intelligence Labs.

Margaret Zhang, CEO at ThirdBrain Labs.

#reading-club-13-0620: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620

#reading-club-12-0613: Origami Robotics (YC W26) on Dexterity

#cvpr-denver-11-0606: 🤖🥘 Saturday Robotics x Manycore Tech x Neural Motion | CVPR 2026 Denver Research Night | Robotics & World Models Reading Club 11

Junfan Zhu & Aurora Feng, Founders of Saturday Robotics

Anthony Zhao, Head of North America at Manycore Tech SpacialVerse

Aurora Feng, Founder at Neural Motion. NM-GenET.

Max Zhaoshuo Li, Robotics and World Model Tech Lead at NVIDIA Cosmos. Cosmos 3.

Xiaofan Li, World Model Tech Lead at X Square Robot. WALL-WM.

Zesen Zhao, University of Michigan. Test-Time Scaling for World Action Models via Zero-Shot Geometric Verification.

Pengyi Liao. VGGT-Ω: From 3D Reconstruction to Scalable Spatial Representation.

Jie Wang, University of Pennsylvania, GRASP Lab. Toward a Robotics MMLU: Lessons from Sim & Real Evaluations of Generalist Policies.

Gordon Qian, Senior AI Researcher at Snap. Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning.

#reading-club-10-0530: Bringing Robots to Life — Learning Humanoid Instincts from the Body Up | San Francisco 0530

Haochen Shi (Stanford, co-advised by Karen Liu & Shuran Song)

#private-dinner-01-0529: Robo Plov x Saturday Robotics

#reading-club-09-0523: CVPR Warm-up & Founders Spotlight — DeltaWorld + VisuoTactile Dexterous Hands

Tommie Kerssies (Amazon Frontier AI & Robotics)

Arjun Subramaniam (Factory Intelligence)

#reading-club-08-0516: Embodied Human Data as the “Internet of Motion and Behavior”

Ryan Punamiya (NVIDIA Gear, Georgia Tech)

#reading-club-07-0509: Learning to Dream: World Models, Imagination, Path to Foundation Models for Control

Ahmet Şemi ASARKAYA (Agility Robotics)

#reading-club-06-0502: Evolution of Video World Models for Robotics

Tongzhou Mu (Rhoda AI)

#reading-club-05-0425: World Models for Physical Intelligence: From Predictive Brains to Embodied Robots

Daniel Dugas & Sergio Arnaud (Meta FAIR)

#reading-club-04-0418: Abstractions of the Physical World for Decision-Making

Siming He (UC Berkeley)

#reading-club-03-0411: Robotic Policy Adaptation

Haoyi Niu (UC Berkeley)

#reading-club-02-0404: JEPA Zoo

Julian Saks (JulianSaks)

#reading-club-01-0328

Join Discord Community

Join Discord Server

Follow Saturday Robotics

saturdayrobotic

Follow YouTube

saturdayrobotic

Subscribe to Luma Calendar

https://luma.com/saturdayrobotic

Location
Please register to see the exact location of this event.
San Francisco, CA
Avatar for Saturday Robotics
Presented by
Saturday Robotics
🤖 Saturday Reading Club on Robotics & World Models for AI Researchers in SF
Hosts: Junfan Zhu, Aurora Feng
discord.gg/WH7DrTHRXK
84 Went