

Serving New LLM Architectures in Production
Registration
Tickets
1
About Event
Modern LLM architectures are evolving fast, but turning new ideas into reliable, high-performance production serving remains a systems challenge.
Join vLLM × Ant Ling for a technical conversation on what it takes to support and optimize emerging LLM architectures in practice. From runtime abstractions and memory management to model integration, profiling, and accelerator-aware optimization, we’ll explore how engineering teams move from initial architecture support to real-world production performance.
Whether you build inference infrastructure, develop models, or optimize AI systems at scale, this session will offer practical insights into bringing next-generation LLM architectures into production.