

Researcher Night: Grace Gong x Patronus AI
Join us for an evening at the frontier of AI research, hosted by Patronus AI and Grace Gong at Patronus HQ in SF!
Hear from researchers pushing the boundaries of post-training and evaluation through a series of fast-paced spotlight talks.
The talks
Yoshi Fujinuma on SpeedrunBench: a benchmark that asks not whether an agent can finish a game, but how fast.
Mariya Vasileva on Beliefload: an evaluation of how LLMs formulate and revise hypotheses in light of new evidence, and how they diagnose and repair a broken environment without overwriting the parts that already work.
Anmol Gulati on Beyond Rows to Reasoning: an agentic framework for reasoning over and editing enterprise spreadsheets with millions of cells, cross-sheet dependencies, and embedded charts.
Nick Saban on GlobeBench: a benchmark for whether language models can faithfully simulate the environments we train agents in.
After the talks, stick around for happy hour drinks, sushi, merch and connect with fellow researchers, engineers, and builders working on some of the most interesting problems in AI today.