Sep 4: Thinking with Video — Video Generation as a Multimodal Reasoning Paradigm by Jingqi Tong
Every week we pick one paper and go deep — video generation, world models, physical reasoning, diffusion, flow matching, and everything in between.
Abstract: "Thinking with Text" and "Thinking with Images" improve LLM/VLM reasoning, but images capture only single moments and keep text and vision separate. "Thinking with Video" leverages video generation models such as Sora-2 to use video frames as a unified medium for multimodal reasoning. The accompanying Video Thinking Benchmark (VideoThinkBench) covers vision-centric tasks (e.g., eyeballing puzzles) and text-centric tasks (e.g., GSM8K, MMMU). Sora-2 emerges as a capable reasoner — comparable to SOTA VLMs on vision-centric tasks (surpassing GPT-5 by 10% on eyeballing puzzles), 92% on MATH and 69.2% on MMMU.
Paper (CVPR 2026): arXiv:2511.04570
Speaker: Jingqi Tong — Researcher at Fudan University (OpenMOSS team, Xipeng Qiu's group), working on multimodal reasoning and video generation.
Website: https://journal.video-reason.com/ Register on this page to receive the Zoom join link (sent automatically with your confirmation).
Subscribe to our mailing list for weekly talk announcements: https://forms.gle/ebgyvtLRz8ABTfdX6