Cover Image for Reading Group: OSWorld 2.0
Cover Image for Reading Group: OSWorld 2.0
Avatar for Snorkel AI Community Events

Reading Group: OSWorld 2.0

Zoom
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

Join this first-ever virtual Snorkel AI Reading Group, a recurring forum to explore the latest frontier developments in AI while building meaningful connections within the community.

In this session, Mengqi Yuan (XLANG Lab, University of Hong Kong) will present OSWorld 2.0.

Among other things, you'll learn:

  • Why existing computer-use benchmarks fail to capture the realism and complexity of real professional workflows, and how OSWorld 2.0's 108 tasks across 31 self-hosted websites close that gap.

  • How partial-credit scoring (27.25 checkpoints per task on average) reveals that frontier agents make real progress but rarely finish, with the best system completing only 20.6% of tasks end to end.

  • Why task horizon acts as a hard limit: binary completion falls toward zero once workflows stretch past roughly 160 minutes, regardless of model.

  • The ten challenge phenomena the benchmark isolates, including cross-source reasoning, implicit-state inference, and dynamic environments, and why agents tend to guess rather than ask the user when state is hidden.

  • Concrete failure traces (a missed purchase approval, a mismatched CAD model, a reimbursement claim gathered but never submitted) showing exactly where long-horizon agents lose the thread.

Paper and results are available on the OSWorld 2.0 project page and the Snorkel AI leaderboard.

Avatar for Snorkel AI Community Events