

Bay Area Frontier Research Club #7 | Google Ventures (dinner + paper discussion)
Scaling Laws, Cognitive Behaviors, Verifier-Hard Evaluation & Self-Improving Agents. Four Frontier Research Talks + Rigorous Q&A.
Four frontier research talks on one stage — Bonnie Li (Google DeepMind) on The Art of Scaling Reinforcement Learning Compute for LLMs (400,000+ GPU-hours, predictive scaling laws for RL, ScaleRL), Kanishk Gandhi (Stanford CS) on the cognitive behaviors that separate the LLMs that self-improve under reinforcement learning from the ones that plateau, Erica Zhang (Stanford PhD / Jump Trading AI Fellow) on TERMS-Bench — what happens when the domain isn't verifiable — and Vignesh Baskaran (Hexo Labs) with the first public preview of SIA, Hexo's self-improving agent framework, ahead of its open-source release the following week.
The Bay Area Frontier Research Club is a curated forum for rigorous discussion on how AI is reshaping the scientific research process. We convene experimental researchers, computational scientists, and research engineers across domains to examine concrete work—papers, methods, and workflows—covering literature synthesis, hypothesis generation, experimental design, simulation, analysis, and reproducibility.
For each session, we curate 2–4 papers selected for rigor and discussion value. Presentations are intentionally brief so the majority of time is reserved for questions and critique: assumptions, evaluation methodology, failure modes, and what would constitute convincing evidence. Papers and supporting materials are shared in advance to ensure a high-baseline conversation.
Agenda
5:30pm: Doors open
5:30pm – 6:30pm: Networking + snacks
6:30pm – 8:00pm: Research presentations + discussion
8:00pm – 8:30pm: Networking
Presenters & topics
Talk #1: The Art of Scaling Reinforcement Learning Compute for LLMs
Bonnie Li is an AI Researcher at Google DeepMind, where she focuses on pushing the boundaries of frontier AI models and agentic post-training.
Her work is at the frontier of developing foundation models, world models, and generalist agents — with core contributions to Gemini 2.5, Gemini 3, SIMA 2, and Genie 2. Previously she worked on impactful research in reinforcement learning, presenting at top-tier ML conferences.
Her paper studies how RL performance scales with compute in large language models. Across experiments costing more than 400,000 GPU-hours, the work shows something quietly provocative: many of the RL training techniques the field has been celebrating mainly improve efficiency, not final capability. From there, the paper introduces scaling laws that predict model performance from smaller runs, and proposes ScaleRL — a practical training recipe demonstrating stable, predictable RL scaling for reasoning-focused LLMs.
Talk #2: Cognitive Behaviors That Enable Self-Improving Reasoners
Kanishk is a PhD researcher at Stanford CS working at the intersection of reasoning, cognition, and reinforcement learning.
His paper (already one of the most-discussed RL-reasoning results of the year) answers a question that has been quietly bothering everyone in the field: Why does Qwen-2.5-3B dramatically self-improve under reinforcement learning, while Llama-3.2-3B — same size, same training compute, same RL recipe — barely moves?
Kanishk's team identified four cognitive behaviors that separate the reasoners that scale from the ones that plateau: verification, backtracking, subgoal setting, and backward chaining. The talk walks through the experimental design, what changes when you prime models with these behaviors, and what it implies for the next generation of reasoning systems.
REVIEW THE PAPER HERE
REVIEW CODE AND PROMPTS HERE
Talk #3: TERMS-Bench: Evaluating LLMs in Semi-Verifiable Domains
Erica is a Stanford PhD, Jump Trading AI Research Fellow, and former Amazon Applied Scientist.
Her latest work, TERMS-Bench (Testbed for Economic Reasoning in Multi-turn Strategy), closes the arc by taking on the hardest evaluation problem in RL-for-LLMs: what happens when the domain isn't verifiable? Math has a right answer. Code passes a test. But most of the work we actually want frontier models to do — negotiation, strategy, judgment, multi-turn reasoning under partial information — does not.
TERMS-Bench is a Bayesian-game framework that makes the environment itself the verifier — by specifying the counterpart's latent type, policy, and payoff structure — instantiated in bilateral price negotiation. The benchmark evaluates 15 frontier models across four orthogonal diagnostic axes, exposing structured failure modes that outcome-only metrics like deal rate hide entirely.
Talk #4: SIA: Self-Improving Agents — A First Look
Vignesh is co-founder and CTO of Hexo Labs, building the research-infrastructure layer for the next generation of agent systems.
Tonight, he is giving the room the first public preview of SIA — Hexo's self-improving agent framework — ahead of its open-source release the following week.
The throughline of the night lands here: the three preceding talks describe how reinforcement learning, cognitive behaviors, and semi-verifiable evaluation come together at the frontier. SIA is what it looks like when you actually ship it into a production agent that gets better over time.
Want to present your work?
If you have a research paper you’d like to discuss with a cross-disciplinary room, submit it for consideration.
SUBMIT YOUR PAPER HERE.
Who should attend
Experimental researchers
Computational scientists across domains (bio/chem/materials/climate/neuro/physics)
Research engineers + lab automation people
Those building tools for literature review, experiment planning, robotics, simulation, or scientific data
No ML background required. If you’ve ever wished research moved faster, you belong here.
Capacity is limited.
We will take photos and short video clips for event recap and promotion. By attending, you consent to being photographed and recorded, and to the use of those images and clips by the organizers on social media and other event marketing channels.
Hosted by
Frontier Syndicate is a venture community connecting frontier tech researchers, builders, and investors through curated convenings and early-stage capital. Across the Bay Area, we host a recurring series of research forums, builder nights, and intimate investor dinners — and back exceptional companies emerging from the labs, communities, and technical networks we convene.
To learn more, contact kristopher@frontiersyndicate.vc
Hexo Labs is a neolab for recursive self intelligence, building open agent systems that help scientific discovery take shape.
Hexo is working to make scientific discovery faster, more inspectable, and more capable of translating breakthrough ideas into real-world impact.
On May 21, Hexo will release SIA, its open-source self-improving agent framework: a system for agents that learn from experiments, evaluate their own progress, and refine their methods over time.
Special thanks to our venue partner: Google Ventures
Google Ventures (GV) is a venture capital firm investing in category-defining companies, with deep roots across AI, biotech, and the broader innovation ecosystem. GV supports founders and researchers building frontier technologies—and helps connect high-caliber builders, operators, and scientific talent through its network and community.