Cover Image for Robotics & WM Reading Club 30: What Scale Doesn't Fix: Representation & Architecture in Robot Manipulation+Omni-Reason Video Model. SF 9/26
Cover Image for Robotics & WM Reading Club 30: What Scale Doesn't Fix: Representation & Architecture in Robot Manipulation+Omni-Reason Video Model. SF 9/26
Avatar for Saturday Robotics
Presented by
Saturday Robotics
🤖 Saturday Reading Club on Robotics & World Models for AI Researchers in SF
Hosts: Junfan Zhu, Aurora Feng
discord.gg/WH7DrTHRXK
138 Went

Robotics & WM Reading Club 30: What Scale Doesn't Fix: Representation & Architecture in Robot Manipulation+Omni-Reason Video Model. SF 9/26

Register to See Address
San Francisco, CA
Registration
Past Event
Please click on the button below to join the waitlist. You will be notified if additional spots become available.
About Event

​Robotics & World Models Reading Club 30: What Scale Does Not Fix: Representation and Architecture in Robot Manipulation. SF 9/26

​A high-signal reading group for AI researchers & builders pushing the frontiers of robotic world models, WAMs, and embodied intelligence. In our previous sessions, we brought together researchers and engineers from Boston Dynamics, Google DeepMind, NVIDIA, Stanford, UC Berkeley, Physical Intelligence, Tesla, Generalist, Rhoda AI, and leading Bay Area robotics startups.

​Hosted by Junfan Zhu & Aurora Feng.

​Support Saturday Robotics Inc: https://donate.stripe.com/28EcN52rjgeY1fJboYgEg00

​

​​​​​Reading Club 30's Core Theme

​What Scale Does Not Fix: Representation and Architecture in Robot Manipulation

​Keynote: Devol Robots (Elijah Yi Herng Ong - CTO)

​Robot foundation models have scaled rapidly over the past two years, yet scale alone has not produced reliable physical interaction. This talk presents two architectures from Devol Robots that attack the problem from opposite ends: what a model represents, and how that model is organized.

​The first is the Interaction World Model (IWM 1.0), which addresses representation. Built on Interaction Geometry, a Riemannian framework that keeps robot sensorimotor signals on the manifolds where they physically live rather than flattening them into Euclidean vectors, IWM infers compliance from the shape of a demonstration itself and commands the robot through an impedance controller whose stiffness and damping the model predicts, with no force or torque signals anywhere in training. On contact-rich tasks from socket insertion to sub-millimeter lens seating it reaches 92% average success against 38% and 50% for leading VLA and world model baselines, using roughly an order of magnitude fewer action parameters.

​The second is Devol-ONE, which addresses architecture. Existing world-action models keep prediction and policy structurally separate and connect them only through a predicted output. Devol-ONE instead runs vision-language reasoning, latent world prediction, and action generation as three streams of a single autoregressive Mixture of Transformers that exchange information at every layer, so imagination shapes execution continuously rather than at one bottleneck. It reaches 98.4% on LIBERO and 71.4% zero-shot on LIBERO-Plus without training on the perturbed data, and holds its performance on RoboTwin 2.0 under environment randomization while a strong pixel-space baseline collapses from 76.6% to 2.7%.

​Elijah will show how these two lines converge: a representation that captures the physics of contact, carried by an architecture that reasons about the future while it acts. Together they define what a deployable and reliable manipulation system looks like in industrial settings, and he will demonstrate what each looks like running on production hardware today.


​Keynote 2

​Omni-Reason: A General Video Model That Could Solve 100 Computer Vision Tasks

​Hokin Deng, Research Scientist @ Harvard, CMU

​Emergence of reasoning through scaling laws has hallmarked artificial intelligence. Before language reasoning, there has been 1000 tasks in the field of natural language processing. But now, every task becomes a sub-task of language reasoning. In the same way, we see the arrival of a generalized video reasoning model that could solve many computer vision tasks as an important milestone in the scaling law of vision models. In this belief, we present Omni-Reason. First, Omni-Reason-Data, 9.33 millions samples training data corpus with an end-to-end data infrastructure that scales mid-training data to at least 10,000 samples per task across 224 computer vision tasks. Second, Omni-Reason-Bench, a carefully curated video-to-video benchmark for assessing emergent reasoning in video models. Third, Omni-Reason-1, the first generalized video reasoning model to achieve above-chance performance on all 100 tasks. Omni-Reason-1 is trained on Amazon Trainium, for which we built a training stack from scratch that is both fast and compute-efficient. We further report a non-trivial trade-off between data mixtures. We open-source the entire pipeline, data infrastructure, training data, benchmark, training stack, and training recipes, to accelerate the arrival of general visual reasoning models.


​​​​​​​​​​​​​​Pre-Readings

​www.devolrobots.ai/whitepaper


​​​​​​​​​​​​Location

​SF

​​​​​​​​​​​​​​​Date & Time

​Saturday, September 26, 2026 | 2:00 PM – 5:00 PM

​​​​​​​​​​​​​​​​Join Discord Community

​Join Discord Server

​​​​​​​​​​​​​​​Follow Saturday Robotics on X

​/saturdayrobotic


​​​​​​​​​​​​​​​Agenda

​2:00 PM – 2:30 PM Door Opens & Social

  • ​Food 😋, beverages🧋 and UNLIMITED strawberries 🍓 (our official reading club fruits ☺️😄).

​2:30 PM – 3:30 PM Keynote by Devol Robots, Hokin Deng

​YouTube Recording: TBD (We are looking for recording volunteers)

​3:30 PM – 5:00 PM Q&A, ​open-floor roundtable (10–20 min per topic) on spotlight papers or any paper you’d like to highlight. Feel free to share why the paper matters and its technical details.


​​Future events

​#iros-reading-club-31-0928: 🍾 IROS 2026 x Saturday Robotics — Robotics Research Night | Reading Club 31. Pittsburgh 9/28


​​​​​​​​​​​​​​​​​​​​​Past events

​#reading-club-28-0912: Booster T2: The Next Frontier of Open Humanoid Robotics

​#reading-club-27-0905: Rethinking Robot Development: Co-Designing Morphology, Sensing, and Learning

​Chenyang Ma (Applied Intuition, UNC)

​#reading-club-26-0829: Sunnyvale 8/29

​#reading-club-25-0822: Contact-Rich Robot Learning from Human Videos and Tactile. SF 8/22

​Kelin Yu (Maryland, Amazon FAR)

​#reading-club-24-0820: Saturday Robotics x EAI @ WRC Beijing | PHYSICAL AI UNPLUGGED —— 与 Physical AI 研究者的一场非正式夜谈

​• 李晓帆 — Head of World Models, X Square Robot (自变量机器人)

​• Tongzhou Mu — Research Scientist, Rhoda AI

​• Haoyi Niu — University of California, Berkeley

​• Yufei Wang — Member of Technical Staff, Genesis AI

​• Shibo Zhao — Research Scientist, Amazon FAR

​• 郑思鹏 — Technical Partner, Being Beyond

​• Tianle Zhang — Embodied AI Researcher, JD Explore Academy

​• Yihang Li — Embodied AI Researcher, JD Explore Academy

​#reading-club-23-0815: Engineering Robotic Simulators for Evaluation and Beyond. SF 8/15

​Kaifeng Zhang (Columbia, World Labs)

​#reading-club-22-0808: ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow. SF 8/8

​Haoyi Niu (UC Berkeley)

​#reading-club-21-0801: Vision-Language-Kinematics Supervision for Perception-Based Humanoid Loco-Manipulation — SF 8/1

​Yen-Jen Wang (UC Berkeley, Amazon FAR)

​#reading-club-20-0725: Agentic Robotics Models (ENPIRE and Cap-X), Mountain View 07/25

​Haoru Xue (UC Berkeley)

​Kris Hauser (Samsung Research America, Robot Intelligence Lab)

​#SIGGRAPH-reading-club-19-0722: SIGGRAPH x Saturday Robotics — World Models for Robotics: Bridging Graphics, Simulation & Physical Intelligence | Reading Club 19, LA 07/22

​#reading-club-18-0718: Causal World Models For Real-World Intelligence. SF 07/18

​Guanming Wang & Bill (General Instinct, YC P26)

​Feng Fan (UCSD, Aether AI)

​#reading-club-17-0711: Soft Tactile-Centric Multimodal Intelligence Toward Safe and Dexterous Manipulation. SF 07/11

​Quan Luu, Purdue.

​#private-lunch-icml-0709: Saturday Robotics x ICML Private Lunch (Seoul)

​#reading-club-16-0704: The Embodied AI Hardware Stack — Supply Chain, Sensors, and the Data Flywheel — SF 07/04

​Jerry Huang, Robotics Center of Silicon Valley.

​#reading-club-15-0627: Scaling Touch: Flexible Tactile Skin for Dexterous Manipulation

​Binghao Huang, Columbia, Amazon FAR.

​#deep-tech-week-14-0625: Deep Tech Week Research Night

​SPEAR: A Simulator for Photorealistic Embodied AI Research. (ECCV 2026 accepted) by Mike Roberts, Senior Research Scientist, Adobe Research.

​Bogdan Cristei, Venture Partner at SHACK15 Ventures.

​Simone Totaro, CTO at Saturn Dynamics.

​Shumo Chu, CEO at General Intelligence Labs.

​Margaret Zhang, CEO at ThirdBrain Labs.

​#reading-club-13-0620: HumanEgo: Train Robot Policy from 30 min Egocentric Videos — SF 0620

​#reading-club-12-0613: Origami Robotics (YC W26) on Dexterity

​#cvpr-denver-11-0606: 🤖🥘 Saturday Robotics x Manycore Tech x Neural Motion | CVPR 2026 Denver Research Night | Robotics & World Models Reading Club 11

​Junfan Zhu & Aurora Feng, Founders of Saturday Robotics

​Anthony Zhao, Head of North America at Manycore Tech SpacialVerse

​Aurora Feng, Founder at Neural Motion. NM-GenET.

​Max Zhaoshuo Li, Robotics and World Model Tech Lead at NVIDIA Cosmos. Cosmos 3.

​Xiaofan Li, World Model Tech Lead at X Square Robot. WALL-WM.

​Zesen Zhao, University of Michigan. Test-Time Scaling for World Action Models via Zero-Shot Geometric Verification.

​Pengyi Liao. VGGT-Ω: From 3D Reconstruction to Scalable Spatial Representation.

​Jie Wang, University of Pennsylvania, GRASP Lab. Toward a Robotics MMLU: Lessons from Sim & Real Evaluations of Generalist Policies.

​Gordon Qian, Senior AI Researcher at Snap. Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning.

​#reading-club-10-0530: Bringing Robots to Life — Learning Humanoid Instincts from the Body Up | San Francisco 0530

​Haochen Shi (Stanford, co-advised by Karen Liu & Shuran Song)

​#private-dinner-01-0529: Robo Plov x Saturday Robotics

​#reading-club-09-0523: CVPR Warm-up & Founders Spotlight — DeltaWorld + VisuoTactile Dexterous Hands

​Tommie Kerssies (Amazon Frontier AI & Robotics)

​Arjun Subramaniam (Factory Intelligence)

​#reading-club-08-0516: Embodied Human Data as the “Internet of Motion and Behavior”

​Ryan Punamiya (NVIDIA Gear, Georgia Tech)

​#reading-club-07-0509: Learning to Dream: World Models, Imagination, Path to Foundation Models for Control

​Ahmet Şemi ASARKAYA (Agility Robotics)

​#reading-club-06-0502: Evolution of Video World Models for Robotics

​Tongzhou Mu (Rhoda AI)

​#reading-club-05-0425: World Models for Physical Intelligence: From Predictive Brains to Embodied Robots

​Daniel Dugas & Sergio Arnaud (Meta FAIR)

​#reading-club-04-0418: Abstractions of the Physical World for Decision-Making

​Siming He (UC Berkeley)

​#reading-club-03-0411: Robotic Policy Adaptation

​Haoyi Niu (UC Berkeley)

​#reading-club-02-0404: JEPA Zoo

​Julian Saks (/JulianSaks)

​#reading-club-01-0328

​Join Discord Community

​Join Discord Server

​Follow Saturday Robotics

​/saturdayrobotic

​Follow YouTube

​/saturdayrobotic

​Subscribe to Luma Calendar

​https://luma.com/saturdayrobotic

Location
Please register to see the exact location of this event.
San Francisco, CA
Avatar for Saturday Robotics
Presented by
Saturday Robotics
🤖 Saturday Reading Club on Robotics & World Models for AI Researchers in SF
Hosts: Junfan Zhu, Aurora Feng
discord.gg/WH7DrTHRXK
138 Went