Cover Image for Sep 4: Thinking with Video — Video Generation as a Multimodal Reasoning Paradigm by Jingqi Tong
Cover Image for Sep 4: Thinking with Video — Video Generation as a Multimodal Reasoning Paradigm by Jingqi Tong
Avatar for Video Model Journal Club
Hosted By
3 Going

Sep 4: Thinking with Video — Video Generation as a Multimodal Reasoning Paradigm by Jingqi Tong

Zoom
Registration
Welcome! To join the event, please register below.
About Event

Every week we pick one paper and go deep — video generation, world models, physical reasoning, diffusion, flow matching, and everything in between.

Abstract: "Thinking with Text" and "Thinking with Images" improve LLM/VLM reasoning, but images capture only single moments and keep text and vision separate. "Thinking with Video" leverages video generation models such as Sora-2 to use video frames as a unified medium for multimodal reasoning. The accompanying Video Thinking Benchmark (VideoThinkBench) covers vision-centric tasks (e.g., eyeballing puzzles) and text-centric tasks (e.g., GSM8K, MMMU). Sora-2 emerges as a capable reasoner — comparable to SOTA VLMs on vision-centric tasks (surpassing GPT-5 by 10% on eyeballing puzzles), 92% on MATH and 69.2% on MMMU.

Paper (CVPR 2026): arXiv:2511.04570

Speaker: Jingqi Tong — Researcher at Fudan University (OpenMOSS team, Xipeng Qiu's group), working on multimodal reasoning and video generation.

Website: https://journal.video-reason.com/ Register on this page to receive the Zoom join link (sent automatically with your confirmation).

Subscribe to our mailing list for weekly talk announcements: https://forms.gle/ebgyvtLRz8ABTfdX6

Avatar for Video Model Journal Club
Hosted By
3 Going