

Beyond Benchmarks: Building an Intelligent AI Router for Production.
Most AI discussions ask "which model is best?" Production systems need to ask a different question: "which model should handle this request?"
In this session, we build a production-style AI inference gateway live, from scratch one that intelligently routes requests between frontier models (GPT, Claude) and locally hosted models (Llama, Qwen) based on data sensitivity, reasoning complexity, latency, context length, hardware availability, and cost. We'll compare identical requests across local and frontier models while visualizing real-time metrics on latency, throughput, cost, and routing decisions.
Who this is for: Data engineers, ML/AI engineers, and architects building or evaluating production AI systems not a benchmark comparison talk, and not for those looking for a single "best model" answer.
What you'll leave with:
How to design an AI inference gateway supporting multiple model providers
When local models are the better engineering call vs. when frontier models justify their cost
Trade-offs across cost, latency, privacy, hardware, fine-tuning flexibility, and quality
Architectural patterns for observable, resilient, vendor-agnostic AI systems