Cover Image for Why Frontier Reasoning Models fail on Interactive 2D Mazes
Cover Image for Why Frontier Reasoning Models fail on Interactive 2D Mazes
Avatar for Manifold Research
Presented by
Manifold Research
We're a mission driven R&D institute dedicated to advancing fundamental discoveries and carrying them to real-world impact.

Why Frontier Reasoning Models fail on Interactive 2D Mazes

Zoom
Registration
Welcome! To join the event, please register below.
About Event

Join us for a live research talk on an early preview into MultiNet v2.0 - our cross-domain, multimodal, long-horizon agentic benchmark.

Vision-language models are increasingly expected to operate as agents that perceive, reason, and act across extended sequences of interactions in dynamic environments. However, state of the art models today perform poorly on long-horizon workflows and fail in a myriad of ways which are not captured by singular scores that agentic benchmarks report. In this early preview to our benchmark we design and build controllable, interactive 2D maze environments to isolate and attribute these failures. Mazes are highly simplified proxies for real-world software workflows - they require agents to take actions over several steps in a constantly changing environment, encounter novel scenarios, reason about how to operate the components of the environment, and make progress towards a goal. By systematically varying environmental complexity and examining behavior across planning, action execution, error recovery, visual association, and causal reasoning, we can identify where long-horizon performance begins to break down. The goal is to move beyond asking whether a model succeeds and toward understanding how and why it fails, providing a more diagnostic foundation for building multimodal agents that can act reliably over time. We were surprised to see that frontier models struggled on simple 2D mazes - and more surprised by how differently each model failed. Tune in to dive deeper into our findings.

The talk will be followed by an open Q&A and discussion.

The talk will be presented by Sean Rivera, an Open Source Research Scientist at Manifold Research working on open-source tool-use action models and multimodal action benchmarks. Sean previously worked as a Senior Embedded Systems Security Engineer at CENSUS and a Security Consultant at IOActive, and completed his Ph.D. and postdoctoral research at the University of Luxembourg on robotic systems and embedded security. His work has spanned robotic security, ARM and trusted execution environments, fuzzing, distributed systems, software-defined networking, eBPF, FPGA development, GPS, and signal processing, with earlier research and engineering experience at CU Boulder and MITRE.

Avatar for Manifold Research
Presented by
Manifold Research
We're a mission driven R&D institute dedicated to advancing fundamental discoveries and carrying them to real-world impact.