Inference Office Hours with SGLang: Performance Optimizations for LLM Serving
Registration
Past Event
About Event
Join us to find out the latest inference optimizations for leading open source models from SGLang on NVIDIA GPUs.
We’ll dive into an in-depth overview of the engineering behind optimizing SGLang on GB200, followed by a deep dive of recent performance advancements targeting low-latency inference and throughput efficiency.
