Cover Image for 90/30 Club (ML reading) #48: Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Cover Image for 90/30 Club (ML reading) #48: Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Avatar for 90/30 Club
Presented by
90/30 Club
We meet weekly in-person to talk about new ML papers! Come and join the discussion!
68 Went

90/30 Club (ML reading) #48: Inside vLLM: Anatomy of a High-Throughput LLM Inference System

Register to See Address
San Francisco, California
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Week 48: Inside vLLM: Anatomy of a High-Throughput LLM Inference System

Paper Link

Modern LLM progress isn’t bottlenecked by model quality anymore, it’s bottlenecked by serving. This deep dive breaks down how systems like vLLM actually achieve high-throughput, low-latency inference at scale, from kernel-level memory tricks to distributed serving architecture.

This post builds a full-stack mental model of inference: starting from a single-GPU offline engine and scaling all the way to multi-node production systems, with a focus on the mechanisms that actually matter in practice—scheduling, KV-cache management, and decoding strategies.


Join us at Mox to explore:

  • How paged attention reframes KV-cache as a memory virtualization problem (and why it’s the core unlock for throughput)

  • Why continuous batching dominates naive batching—and how schedulers actually decide what runs next

  • Tradeoffs between latency (TTFT, ITL) vs throughput, and how real systems optimize both simultaneously

🔎Analyzed Papers

​Discussion at 20:00, (optional) quiet reading from 19:00.

Location
Please register to see the exact location of this event.
San Francisco, California
Avatar for 90/30 Club
Presented by
90/30 Club
We meet weekly in-person to talk about new ML papers! Come and join the discussion!
68 Went