Frontier Paper Club
Frontier Paper Club — NYC
Papers. People. Possibilities.
How do we know an AI agent is actually getting better? What makes a research result hold up beyond its benchmark? And which ideas are ready to become systems people can depend on?
Frontier Paper Club is a monthly gathering in New York for researchers, engineers, and founders exploring the next generation of AI. Sponsored by Betaworks and Blobfish.ai, we bring together people developing new methods, building agent systems, and testing what works in practice.
Each session centers on 2–3 papers, benchmarks, or technical implementations selected around a shared research question. Topics span AI agents, reinforcement learning, post-training, reasoning, memory, evaluation, and simulation environments.
Presentations are brief and technically grounded, leaving plenty of time to examine the work together: the assumptions behind a method, the strength of its evidence, the details needed to reproduce it, and the failures that suggest where research should go next.
🕒 Evening format
5:30pm Arrivals & introductions
Meet fellow researchers and builders. Share what you’re working on and the questions you’re hoping to explore.
6:00pm Research presentations & benchmark discussion
Presenters walk through the problem, approach, key results, and limitations. Each talk is followed by audience questions and technical discussion.
Presenter 1: John Cai - Reflection AI
https://arxiv.org/abs/2608.13167
As VLMs become increasingly deployed in the physical world in ambiguous scenarios, it is imperative to ensure that VLMs know when to “not act” instead of picking an action based on insufficient information. We devise a new benchmark to systematically measure how capable VLMs are at abstaining when faced with physical uncertainty. We find a large divergence between VLM capabilities in abstaining when the ambiguity exists in text vs image domains, with textual ambiguity being 4x more readily detectable. Using linear probes, we also show that VLMs have a latent understanding of ambiguity but do not express it externally. We further demonstrate that activation steering could be used to causally induce or reduce abstention, paving the way for safer VLM deployment in the physical world.
8:00pm Open conversation & networking
Continue the discussion, exchange implementation notes, and meet potential collaborators.
The detailed schedule will be announced with the speaker lineup.
📝 Want to present your work?
Have research or an implementation you’d like to discuss with New York’s AI community?
We’re accepting submissions covering:
AI agents, reasoning, planning, and memory
Reinforcement learning and post-training
Benchmarks, evaluation methods, and agent reliability
Simulation environments and synthetic training data
Reproductions, negative results, and practical lessons from deployment
Share a link to your paper, benchmark, project, or slides, along with what you’d like to present. Submissions will be considered for upcoming sessions.
Questions? Contact sam@blobfish.ai.
🤝 Sponsored by
Betaworks builds and invests in technology companies, with a longstanding presence in New York’s startup community and a focus on emerging technologies, including AI, agents, and developer tools. Through its thematic investment and product development program, Camp, it brings founders together to explore new categories and build companies. Its history includes backing companies such as Hugging Face and Stability AI.
Blobfish.ai builds training gyms for AI agents—simulated enterprise environments where models practice workflows across tools such as CRMs, ERPs, and spreadsheets.
Its environments combine realistic data, executable tools, adjustable task difficulty, and deterministic checks of task outcomes. As models improve, Blobfish generates new challenges to support continued learning, with a focus on the complex, multi-step work agents need to perform in practice.
