The One About Agents & Evals (Round II)
Building an agent is one thing. Getting it to production, and knowing you can trust it there, is another.
Join practitioners as they share what it takes to move agents beyond the demo, from the messy realities of building and shipping to evaluating whether they're actually ready for the real world.
More About the Sharings
Nazran (Lead Software Engineer, Zuno Carbon) will share on “Iterating Our Way to Shipping a Data Agent”
Everyone sees the end product. What they don’t see are the failed iterations, unexpected bugs, and tokens burned along the way.
Naz will share Zuno Carbon’s journey building a data agent for its enterprise sustainability platform, from an early proof of concept to a mid-build pivot, flaky evals, and an unexpectedly large token bill. Hear what worked, what broke, and what the team would do differently when taking an agent from a convincing demo to something enterprise users can actually rely on. (Technical Level: 200)
Fahim Surani (APJ Head of Solutions Architecture, Arize) will share on "How to Get Your Agents to Production: Or Die Trying."
Everyone's building AI agents. Almost nobody's shipping them. Fahim has watched countless agents die on the road to prod and he's distilled the survivors' secrets into 10 commandments for getting yours across the line. Expect equal parts hard-won wisdom and questionable advice; telling the two apart is left as an exercise for the audience. Viewer discretion is advised. (Technical Level: 200)
Timothy Lin (Lead Product Manager, Resaro) will share on "Why Evals Don't Scale to Trust"
Passing an eval doesn't necessarily mean an AI system is ready to deploy, especially in high-stakes areas such as healthcare, finance, and security. Drawing on Resaro's experience independently testing AI systems, Tim will explore why this gap becomes even harder to close as systems become more agentic, where conventional evaluation can fall short, and why establishing real-world trust isn't purely an engineering problem. (Technical Level: 200)
More About the Speakers
Khairul Nazran Kamarulnizam is Lead Software Engineer at Zuno Carbon, and one of the rarer extroverted software engineers. Once a serial hackathoner, he now focuses on the harder part: turning creative hacks and POCs into products that can actually scale. His experience spans building products and engineering teams across startups, Xendit, and Bank Negara Malaysia. Naz studied Economics and Computer Science at the University of Michigan, a combination that continues to shape how he thinks about the business and technical sides of building products.
Fahim Surani leads the Solutions Architecture function for APJ at Arize AI, helping enterprises design, evaluate, and operationalize reliable GenAI and agentic AI systems. He brings nearly two decades of experience as a software developer, AI engineer, and architect across industries including telecommunications, gaming, financial services, and cloud technology. Prior to Arize, he was a Senior Solutions Architect at AWS, working with customers on generative AI, machine learning, and scalable cloud architectures. Today, his focus is on agent evaluation, observability, and helping teams move AI systems from experimentation into production.
Tim is Lead Product Manager at Resaro, specialising in synthetic data generation and AI testing infrastructure. Beyond his day-to-day product work, he's an active open source contributor and has helped shape Singapore's AI governance landscape through the AI Verify initiative and MAS Veritas toolkit.
More About The Series
AI Wednesdays is Lorong AI’s weekly gathering, bringing together practitioners, researchers and innovators for technical discussions on research insights, product development and engineering practices.
Get involved: Learn more about Lorong AI | Speaker Sign-up | WhatsApp Community | LinkedIn | X
