Cover Image for Inference Office Hours with SGLang: Performance Optimizations for LLM Serving
Cover Image for Inference Office Hours with SGLang: Performance Optimizations for LLM Serving
Avatar for NVIDIA
Presented by
NVIDIA
Hosted By
2 Went

Inference Office Hours with SGLang: Performance Optimizations for LLM Serving

YouTube
Registration
Past Event
Welcome! To join the event, please register below.
About Event

​Join us to find out the latest inference optimizations for leading open source models from SGLang on NVIDIA GPUs.

​We’ll dive into an in-depth overview of the engineering behind optimizing SGLang on GB200, followed by a deep dive of recent performance advancements targeting low-latency inference and throughput efficiency.

Avatar for NVIDIA
Presented by
NVIDIA
Hosted By
2 Went