Cover Image for handling 10 million agent traces a day: an engineering session
Cover Image for handling 10 million agent traces a day: an engineering session
Avatar for Failproof AI
Presented by
Failproof AI
Auto healing for Agents
1 Going

handling 10 million agent traces a day: an engineering session

Register to See Address
San Francisco, CA
Registration
Welcome! To join the event, please register below.
About Event

​an engineering session on storing and querying agent traces at 10 million a day: the write path, the schema, what it costs to keep, and the designs that fell over first.

​an agent trace isn't a log line. one run can have dozens of nested tool calls, long model outputs, retries and branches, and you need to pull the whole session back in one piece when something goes wrong. at a few thousand runs a day, anything works. at 10 million traces a day, most of the obvious choices break.

​Siddartha A Y, founding engineer at failproof ai, built the trace store and walks through it from the inside: how we store and query agent traces at that volume, and what we changed to get there. this is an engineering session, not a product demo.

​01 · what makes agent traces different
why traces are wider, deeper and more uneven than regular observability data, and what that does to a storage design.

​02 · the write path
how traces come in, how we batch them, and how we keep ingestion from falling behind when traffic spikes.

​03 · schema and layout
how we lay out spans so that "show me this whole session" and "find every failed tool call this week" are both fast.

​04 · keeping cost sane
what we keep hot, what we move to cheaper storage, and where the money actually goes at this volume.

​05 · what broke along the way
the designs we tried first, where they fell over, and what we'd do differently.

​06 · your questions, live
the last 20 minutes are open. bring your own tracing setup and the query that's too slow.

​good for: engineers building agent observability · anyone running agents in production · anyone on high-volume data pipelines

​how the evening runs (all times pdt, subject to change)

  • ​6:00 pm · what agent traces look like at scale (15 min)

  • ​6:15 pm · walkthrough of the storage design, with real numbers: the write path and batching, span layout for whole-session and needle queries, and hot versus cold storage (25 min)

  • ​6:40 pm · open q&a (20 min)

​who's on

  • ​Siddartha A Y, founding engineer @ failproof ai, speaker · x · linkedin

  • ​Nivedit Jain, cto & co-founder @ failproof ai, host · x · linkedin

  • ​Nikita Agarwal, ceo & co-founder @ failproof ai, host · x · linkedin

​questions

​do i need to use failproof to get anything out of this?
no. the session is about storing and querying agent traces at volume, and failproof is the system we hit the problem on. the write path, the span layout and the hot versus cold split are the same questions on any trace store.

​what will i actually see?
the storage design and real numbers, not a slide of a reference architecture.

​can i bring my own problem?
yes. the last 20 minutes are open, starting 6:40 pm pt.

​what is failproof ai?
failproof is your agent's oversight layer, steering it towards success. it traces your agents, finds failure patterns over time and sets up rules to prevent them from happening again. the open-source cli is free: npm i -g failproofai.

​something else? ask on discord.

​tuesday, october 20 · 6:00 to 7:00 pm pdt

Location
Please register to see the exact location of this event.
San Francisco, CA
Avatar for Failproof AI
Presented by
Failproof AI
Auto healing for Agents
1 Going