Cover Image for Paper Discussion: Steering Agents at Runtime
Cover Image for Paper Discussion: Steering Agents at Runtime

Paper Discussion: Steering Agents at Runtime

Hosted by Nivedit Jain & Nikita Agarwal
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

​Steering Agents at Runtime: a paper discussion

​Two hours of nerding out about why agents fail and what you can do about it without touching the model.

​We ran coding agents on T-Bench 2.1 and read every failed run. Most failures weren't the model lacking skill. The agent was missing one fact about its environment: a server that dies when the shell closes, a damaged database opened before it was copied, an output that was right except for one byte. We wrote short runtime policies that catch those moments and tell the agent what it's missing. On tasks where a policy recognized the failure, success went from 37% to 69%. That's the paper: FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents (arXiv:2609.26048). Read online: https://befailproof.ai/research/runtime-policies/

​This is a discussion, not a demo day. Here's what you'll get into:

​Real failed runs, up close. We'll put actual agent traces on screen and walk through where each one went wrong, step by step. The failures are more specific and more fixable than most people expect.

​The policies themselves. Not a diagram of them. The actual code, what triggers each one, and why the exact wording of the instruction mattered more than when it fired.

​The parts we're unsure about. Our text matching recognized only 2 of 18 reworded tasks. One policy set raised cost by almost 48%. One of our comparisons didn't clear significance. We'll bring these to the tables and want your take.

​Your own agent's failures. Bring a failure mode you keep hitting. At the tables, we'll work through whether a runtime policy could catch it, and what it would need to see.

​A panel that's asked to disagree. Researchers who've read the paper in advance will review it in front of the room, and then it's open floor.

​Everything to rerun it. We're releasing the policies, matching rules and task list at the event, so you can try it on your own agents the following week.

​Schedule

  • ​3:00 · chai, find a table

  • ​3:15 · walkthrough of the method, traces and results

  • ​3:35 · table discussions

  • ​4:05 · panel review and open floor

  • ​4:35 · replication kit, hang around and keep talking

​Table discussions follow Chatham House rules. The walkthrough and panel will be recorded.

​Who's in the room: people working on agents, evals or reliability, whether in research or in production. We'll send a short pre-read a week before. You don't need to have read the paper to come, but it'll make the tables better.

​Saturday, October 24 · 3:00 to 5:00 pm IST · HSR Layout, Bengaluru

Location
1st Sector
HSR Layout, Bengaluru, Karnataka 560102, India