

Frontier Tower AI Paper Reading Club - Week 20 - Where Do Deep-Research Agents Go Wrong?
Weekly paper:
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
https://arxiv.org/abs/2606.02060
Abstract:
"Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not which parts of the trajectory make the answer unreliable. We study span-level error localization for deep-research agents. We collect 2,790 real trajectories from two agent frameworks, three backbone models, and three benchmarks, convert raw logs into semantic spans, and annotate harmful error spans through LLM-assisted expert review. From these annotations, we build TELBench, a 1,000-instance benchmark for identifying error spans among normal exploration, failed searches, tentative hypotheses, and harmless noise. We further propose DRIFT, a claim-centric auditing framework that tracks agent claims, checks their support in trajectory evidence, and marks spans where unsupported or conflicting claims affect the answer path. Experiments across model families and auditing frameworks show that DRIFT improves span-level error localization and first-error accuracy by up to 30 percentage points. Our work provides a process-level view of reliability in deep-research agents."
What are the group goals? Stay on top of AI research, improve understanding of AI fundamentals+math.
Who is welcome? Everyone! Try to put in at least some time on the paper and come prepared with questions or things you'd like to discuss, but it's ok to just show up.
How will it work? Ideally everyone shows up and has at least made one solid pass through the paper, and then we can spend 30-60 minutes as a group going over it.
This event is hosted at the Frontier Tower:
We are transforming a 16-floor tower in San Francisco into a self-governed vertical village—a hub for frontier technologies and creative arts. Tier-one labs presenting AI, Ethereum, biotech, neuroscience, longevity, robotics, makerspace, human flourishing, and arts & music. These floors will house innovators and creators pushing the boundaries of human potential in a post-AI-singularity world.
Apply here for founding citizenship: https://frontiertower.io/apply
Why should I become a citizen?
Be part of creating the first self-governed vertical village
Connect with the most creative people in the city
Get access to all floors, free event space & movement floor
Website: https://frontiertower.io/
Need more reading? Visit https://frontiertower.notion.site/