

Systems Reading Group Featuring vLLM
This one is for the ML infra lovers, just in time for PyTorch conference!
For this session, we'll be joined by Simon Mo (vLLM lead) for what's top of mind for vLLM, a high-throughput and memory-efficient inference and serving engine for LLMs.
Specifically, we'll be discussing the paper "Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding". We will put this paper in context of LLM inference napkin math and how would it fit with vLLM.
https://arxiv.org/abs/2507.19427
For this session, we'll gather ~30-35 folks of different systems backgrounds together to go through the highlights of the paper, answer key questions, and implementation.
We'll have food and drinks provided as well! Please sign up early as the slots filled up very quickly.