

Inference Routing Summit
Where intelligence
meets optimization.
Building model access into production is easy. The hard part is matching each inference call with a model that optimizes token budget, capability and availability, while maintaining agent SLO and avoiding provider lock-in.
Join engineering leaders from Google, Hitachi, Broadcom, Composio, and more at the Inference Routing Summit, hosted by Tetrate, to learn how dynamic routing helps your organization optimize every inference call, at scale. This curated gathering of leaders will feature intimate, off-the-record discussions about deploying AI in production.
In partnership with the Agentic AI Foundation.
What We'll Cover
AI traffic management at scale
MCP and agent connectivity in production
How enterprises are building their inference stack
Speakers
Delyan Raychev, Cloud Infrastructure for AGI / OpenAI
Mandar Jog, Member of Technical Staff / xAI
Alex Salvekar, COO / Agentic AI Foundation (AAIF)
James Waters, VMware / Broadcom
Tina Tsou, Senior VP AI Ecosystem / Bitdeer
Bala Krishnapillai, SVP, CIO & Head of Business incubation/ Hitachi
Anoop Veteth, VP, Product / Google
Soham Ganatra, Co-Founder / Composio
Karthik Kunduru, VP of Engineering / Fiserv
More speakers to be announced.
Why Attend
Hear from enterprises like yours.
Companies running AI in production share learnings from building their inference infrastructure.Focus on real-world problems.
Maintaining agent SLO, token budget optimization, maximizing GPU utilization, avoiding provider lock-in, and more.
Connect & learn with peers.
Curated room of 100 senior leaders responsible for building and maintaining an AI stack in production at scale.
Format
A full day of keynotes, talks, and panels. Lunch included. In person only, no livestream.
Request to Attend
If you're a CTO, VP of Engineering, or head of platform, AI, or infrastructure responsible for building and maintaining an AI stack in production, request access below.
The Venue
Join us at the The SVB Experience Center:
A 120-seat event and networking venue operated by Silicon Valley Bank at 532 Market Street in downtown San Francisco, California.
About Tetrate
Tetrate runs AI traffic for enterprises and service providers, giving them consistent control everywhere their business operates. Its AI gateway technology translates technical AI requirements into policies the business can govern, such as cost, model capability, customer experience, resilience, and sovereignty, then enforces those policies consistently across every gateway and location the business depends on. Teams remain free to choose the models, providers, agent harnesses, guardrails, and other AI tools that fit their needs while the enterprise maintains central control.
For more information, visit www.tetrate.io.