

Daytona AI Builders - SF, September 2026
βAn event dedicated to exploring all things AI Engineering!
βEvent partners: WorkOS & LlamaIndex
ββAgenda
ββββββπ 5:30 pm β 5:35 pm
Welcome and Opening Remarks
βπ€ Borna Perak, GTM & Partnerships @ Daytona
ββββββββπ 5:35 pm β 5:50 pm
Talk "Your Agent's Slowest Tool Is Its Mouse"β
βπ€ Muhammad Annas Hashmi, DevRel at Daytona
ββββββββOutline:
βEvery click a GUI agent makes costs a full screenshot and a model round trip, and it only lands if the layout stayed where the model last saw it. Chain a task out of clicks and you have bought seconds of waiting and a context window full of pixels for work a shell one-liner could have done. Screenshots are heavy. Text is light. And a click sequence is the most fragile program ever written. No variables, no error handling, and the only retry is asking the model again.
This talk builds the cost model of computer use. Where the waiting actually comes from, why pixel agents break when nothing is wrong, and the levers that fix it: read structure (a DOM, an accessibility tree) instead of rendering it to an image first, write code instead of emitting one click at a time, and cache what worked so the second run costs nothing. I'll then follow up with a demo of what we're doing at Daytona to address these.
ββπ 5:50 pm β 6:05 pm
Talk "TBA"
βπ€ Speaker from WorkOS
ββββββββOutline:
βTBA
ββπ 6:05 pm β 6:20 pm
Talk "Why Your Agent Can't Read Your Docs"
βπ€ Yong Park, Member of Technical Staff of LlamaIndex
ββββββββOutline:
βEvery document your agent reads was built for a human eye. A table is a table because of where the lines fall. A column break is obvious because you can see the gutter. Strip the layout out and you get a wall of numbers with no header, two columns spliced into one sentence, and a chart that simply isn't there. The model answers anyway and that's what makes it expensive.
This talk builds the failure map for document ingestion. Where the information actually gets lost, why the same pipeline works on one vendor's invoices and breaks on the next, and how to tell a parsing problem from a retrieval problem before you spend a week tuning the wrong layer. The fixes: preserve structure instead of flattening to text, treat layout as signal, and check the parse output directly instead of judging it through the answers. We'll show exactlyΒ how we're solving this with LlamaParse.
ββπ 6:20 pm β 6:30 pm
Talk "TBA"
βπ€ Speaker TBA
ββββββββOutline:
βTBA
ββπ 6:30 pm β 6:40 pm
Talk "TBA"
βπ€ Speaker TBA
ββββββββOutline:
βTBA
ββββββββββββββββπ 6:40 pm - 8:00 pm
βββNetworking
βWith pizzas and beverages
βAbout event
βThis is dynamic gathering for AI enthusiasts, innovators, and professionals to collaborate, share ideas, and explore the latest advancements in artificial intelligence. Whether you're building AI products, researching cutting-edge algorithms, or simply passionate about the field, join us to connect, learn, and drive the future of AI forward.