Cover Image for Inference Office Hours with SGLang: Performance Optimizations for LLM Serving
Cover Image for Inference Office Hours with SGLang: Performance Optimizations for LLM Serving
Avatar for NVIDIA
Presented by
NVIDIA
Hosted By
2 Went

Inference Office Hours with SGLang: Performance Optimizations for LLM Serving

YouTube
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Join us to find out the latest inference optimizations for leading open source models from SGLang on NVIDIA GPUs.

We’ll dive into an in-depth overview of the engineering behind optimizing SGLang on GB200, followed by a deep dive of recent performance advancements targeting low-latency inference and throughput efficiency.

Avatar for NVIDIA
Presented by
NVIDIA
Hosted By
2 Went