MLn Club (ML Reading Group) #16: Vision Language Action Models: Foundation Models for Robotics
βWelcome to Week 16: Ο0.7: A Steerable Generalist Robotic Foundation Model
βThe Blog: https://www.pi.website/blog/pi07
βThe Paper: https://www.pi.website/download/pi07.pdf
βWhen language models became generalists by training on enormous, diverse datasets, what is the robotics equivalent when useful robot data is scarce, expensive, and fragmented across different machines?
If a robotβs training set contains demonstrations, autonomous rollouts, failures, human video, and Internet data, how does the model learn which behaviors to imitate, which to avoid, and which pieces can be recombined into something new?βThe central problem behind general-purpose robotics is not just collecting more data, but making heterogeneous data usable. Prior vision-language-action models can perform many trained behaviors, yet still struggle to compose skills into new tasks without task-specific fine-tuning. Ο0.7 attacks this by expanding the modelβs prompt beyond a simple language instruction: each trajectory can be conditioned on detailed subtask descriptions, generated subgoal images showing what the near-future world should look like, and episode metadata describing speed, quality, mistakes, and control mode. Architecturally, Ο0.7 is a 5B-parameter VLA with a 4B Gemma 3 vision-language backbone, a memory-style history encoder, and an 860M-parameter flow-matching action expert that predicts chunks of continuous robot actions.
βEach of those context signals is a lever on data diversity: metadata lets Ο0.7 learn from failed and suboptimal autonomous rollouts without blindly imitating them, subgoal images import knowledge from web-scale image-generation pretraining, and cross-embodiment and human data expose the model to behaviors beyond any single robot. The payoff is striking: one model can match task-specific specialists on dexterous tasks, follow unusual instructions in unseen environments, transfer skills between substantially different robots, and compose previously learned behaviors to solve new tasks without task-specific post-training. Ο0.7 asks whether richer context can substitute, at least partly, for pristine robot datasets.
βJoin us at CASI for discussion at 8 pm, and (optional) quiet reading from 7 pm.
βπ Reading Recommendations, Questions, or Comments? Contact us here!
π View past meeting notes here.
βWhat's this?
βA super warm group of folks discussing their favorite topics!
βIn the first half, we host an optional quiet reading space
βIn the second half, we have a discussion where people can talk about what they found interesting about the reading and ask questions about things they didn't understand
βWhen/Where:
βCMU AI Safety Initiative's Office, 201 Craig Street, right across the PNC bank. Look for the open door up the stairs.
β8pm discussion, 7pm optional quiet reading time.
βHere's how it usually goes:
β7:00 PM β arrival and settling in
8:00 PM β introductions
8:10 PM β discussion time
9:00 PM β wrap up then open discussion
βWho's it for?
βPeople who've been wanting to read up on the latest papers in ML and other fields but just haven't been able to find the time/motivation.
βWhy:
βWe've been procrastinating too much on our readings, even though we have so much fun doing them. We know we're not alone in this and want to keep others accountable for learning more about what they're passionate about!
βWe've also met a ton of really fun friends by discussing what we care about!
βRules/guidelines on how to act:
βAct like a host, include people in conversations, talk to people even if they're strangers, offer to explain what you know, and keep an open mind! come to read stuff and find super fun friends :)
βBring snacks if you're feeling kind!
