π€π¨ Sundae Robotics 05: 3D World Models, Spatial Intelligence & Adaptive AI β Modality Forcing & 3D Pre-training Objectives
βπ€π¨ Grab a sundae and join Sundae Robotics, a private, invite-only Sunday series bringing together robotics researchers, founders, and builders working at the frontier of physical intelligence.
βSundae Robotics 05
3D World Models, Spatial Intelligence & Adaptive AI
Featured Talk: Modality Forcing & 3D Pre-training Objectives
βUnderstanding the 3D World: Scalable Pre- and Post-training Recipes for Adaptive AI
βKeynote: Bardienus (Bart) Duisterhof
Final-year PhD Student, Carnegie Mellon University Robotics Institute (Jeffrey Ichnowski) Β· World Labs Β· Collaborator with Deva Ramanan
βHow do we build generative models that deeply understand the 3D world, and adapt quickly to new tasks? In this talk, Bart will present his recent contributions in scalable pre-training and post-training recipes. First, he will present a new pre-training objective for learning 3D dynamics. Inspired by masked representation learning, the approach explores scalable supervision for learning representations that transfer to downstream world modeling and imitation learning.
βSecond, Bart will cover Modality Forcing, a post-training recipe to extract spatial knowledge from image generation models. Modality Forcing trains a single DiT to model the joint distribution between modalities by setting a separate noise level for each modality. The method competes with the very best depth models and scales with text-to-image pre-training.
βFinally, Bart will discuss ongoing work toward more scalable and adaptive AI. On the pre-training side, he will argue that video generation is wasteful, motivating the search for more compressed transition models. For post-training, he will lay out directions toward adaptation across tasks at a fraction of today's cost. Together, these directions aim at a future where anyone can adapt powerful open models to new tasks.
βBart's research develops generative models for spatial intelligence, spanning 3D/4D representation learning and scalable pre- and post-training recipes that transfer to downstream tasks. He is a final-year PhD student at Carnegie Mellon University's Robotics Institute, where he is advised by Jeffrey Ichnowski and collaborates with Deva Ramanan. He currently works on 3D pre-training at World Labs. His recent work includes Modality Forcing, which showed that text-to-image pre-training substantially improves 3D perception. Bart is increasingly interested in the foundations of generation β how these models learn, and how to make them more adaptable and efficient. He believes in a future with powerful open models that can be quickly adapted for new tasks, and is passionate about making open AI beneficial for society.
βPre-Reading
ββ’ Modality Forcing: Extracting spatial knowledge from image generation models
https://modality-forcing.github.io
βTopics
ββ’ Scalable pre-training for 3D world understanding
ββ’ Learning transferable representations of 3D dynamics
ββ’ Extracting spatial knowledge from image generation models
ββ’ 3D/4D representation learning for world modeling and imitation learning
ββ’ Scaling spatial intelligence with text-to-image pre-training
ββ’ Compressed transition models beyond video generation
ββ’ Efficient post-training and adaptation across downstream tasks
ββ’ Open and adaptable foundation models
βOpen Discussion + Q&A
ββ’ What is the right pre-training objective for models that need to understand and act in the 3D world?
ββ’ Can point tracks provide a more scalable representation for world modeling than pixels or video?
ββ’ How much spatial intelligence is already latent inside large image generation models?
ββ’ Is video generation an unnecessarily expensive way to learn world dynamics?
ββ’ How can powerful foundation models be adapted to new tasks at a fraction of today's cost?
