Cover Image for Frontier Paper Club
Cover Image for Frontier Paper Club
62 Went

Frontier Paper Club

Hosted by Samuel & Betaworks
Registration
Past Event
Please click on the button below to join the waitlist. You will be notified if additional spots become available.
About Event

​Frontier Paper Club — NYC

​Papers. People. Possibilities.

​How do we know an AI agent is actually getting better? What makes a research result hold up beyond its benchmark? And which ideas are ready to become systems people can depend on?

​Frontier Paper Club is a monthly gathering in New York for researchers, engineers, and founders exploring the next generation of AI. Sponsored by Betaworks and Blobfish.ai, we bring together people developing new methods, building agent systems, and testing what works in practice.

​Each session centers on 2–3 papers, benchmarks, or technical implementations selected around a shared research question. Topics span AI agents, reinforcement learning, post-training, reasoning, memory, evaluation, and simulation environments.

​Presentations are brief and technically grounded, leaving plenty of time to examine the work together: the assumptions behind a method, the strength of its evidence, the details needed to reproduce it, and the failures that suggest where research should go next.


​🕒 Evening format

​5:30pm Arrivals & introductions
Meet fellow researchers and builders. Share what you’re working on and the questions you’re hoping to explore.

​6:00pm Research presentations & benchmark discussion
Presenters walk through the problem, approach, key results, and limitations. Each talk is followed by audience questions and technical discussion.

https://arxiv.org/abs/2608.13167 As VLMs become increasingly deployed in the physical world in ambiguous scenarios, it is imperative to ensure that VLMs know when to “not act” instead of picking an action based on insufficient information. We devise a new benchmark to systematically measure how capable VLMs are at abstaining when faced with physical uncertainty. We find a large divergence between VLM capabilities in abstaining when the ambiguity exists in text vs image domains, with textual ambiguity being 4x more readily detectable. Using linear probes, we also show that VLMs have a latent understanding of ambiguity but do not express it externally. We further demonstrate that activation steering could be used to causally induce or reduce abstention, paving the way for safer VLM deployment in the physical world.

https://arxiv.org/abs/2606.21804 Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.


https://arxiv.org/abs/2607.28545 Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs a live OpenTelemetry-instrumented microservice system--exposing six days of metrics, logs, and traces through real telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and full source-code access--with 1,079 RCA tasks that systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's κw=0.90). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard--a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access degrades every metric. Crucially, these are performances on a curated 50 GB / six-day testbed with tasks investigated in isolation on a system whose code and instrumentation are public. Since real production systems are order of magnitudes larger, more dynamic, and more idiosyncratic, the gap we report is a lower bound on the engineering investment required before frontier coding agents can be safely entrusted with production reliability.

​

​8:00pm Open conversation & networking
Continue the discussion, exchange implementation notes, and meet potential collaborators.

​The detailed schedule will be announced with the speaker lineup.


​📝 Want to present your work?

​Have research or an implementation you’d like to discuss with New York’s AI community?

​We’re accepting submissions covering:

  • ​AI agents, reasoning, planning, and memory

  • ​Reinforcement learning and post-training

  • ​Benchmarks, evaluation methods, and agent reliability

  • ​Simulation environments and synthetic training data

  • ​Reproductions, negative results, and practical lessons from deployment

​Share a link to your paper, benchmark, project, or slides, along with what you’d like to present. Submissions will be considered for upcoming sessions.

​Submit your work here →

​Questions? Contact sam@blobfish.ai.


​🤝 Sponsored by

​Betaworks

​Betaworks builds and invests in technology companies, with a longstanding presence in New York’s startup community and a focus on emerging technologies, including AI, agents, and developer tools. Through its thematic investment and product development program, Camp, it brings founders together to explore new categories and build companies. Its history includes backing companies such as Hugging Face and Stability AI.

​Blobfish.ai

​Blobfish.ai builds training gyms for AI agents—simulated enterprise environments where models practice workflows across tools such as CRMs, ERPs, and spreadsheets.

​Its environments combine realistic data, executable tools, adjustable task difficulty, and deterministic checks of task outcomes. As models improve, Blobfish generates new challenges to support continued learning, with a focus on the complex, multi-step work agents need to perform in practice.

Location
Betaworks
29 Little W 12th St, New York, NY 10014, USA
62 Went