

AI Scale Talks EP.2 | Accelerating LLM Inference: From Speculative Decoding to Diffusion LLMs
AI Scale Talks is a four-part live webinar series from Lablup. The series theme is "From Cell to Factory": each week we scale up one level, from a single inference engine to full-scale AI infrastructure operations.
EP.2 starts with LLM inference. The response latency and operational cost of LLM services are determined not only by the model itself but also, to a significant extent, by the inference methodology employed. This webinar will begin with an accessible overview of the fundamental architecture of LLM inference, followed by an examination of key optimization techniques, including KV caching, quantization, and speculative decoding. Building on these principles, the session will conclude with a presentation of the architecture and empirical results of our Diffusion LLM model currently in training.
➡️ Format
40-minute session + 10-minute live Q&A
Live on Zoom. The join link is shared with registered guests.
📆 All episodes air Wednesdays at 9:00 AM PT.
EP.1 | Aug 26 | High-performance LLM/VLM inference with MLxcel engine, Jeongkyu Shin
EP.3 | Sep 9 | Deploying an AI Platform on Infrastructure You Don’t Control, Jonghyun Park
EP.4 | Sep 16 | You Built the Token Factory. Who Runs It?, Jongmin Kim