

London Systems Club
London Systems Club is a meetup for engineers interested in systems, hardware, and performance.
We cover topics like operating systems, compilers, storage, networking, GPUs, runtimes, AI infrastructure, and distributed systems through technical talks and discussion.
No sales pitches. No recruitment. Just systems.
Schedule
6:00 – 6:30 PM Arrival & Networking
6:45 – 7:15 PM Kenneth, MTS @ Prime Intellect
Dive into RLM with Prime Agent
Context as a variable: recursive language models & live Prime Agent demo.
Optional pre-reading if you want to get the most out of the talk: Recursive Language Models
7:15 – 7:45 PM Muhammad Hamzah Chariwala, Head of Compute Infra @ Callosum
Characterising Silicon Performance Regimes for Inference: Trainium2
Accelerator performance depends strongly on workload shape, parallelism, data layout, the software stack used to express them, and more. This talk presents a set of initial Trainium2 experiments aimed at identifying favourable and unfavourable characteristics and regimes for LLM inference, using those observations to guide algorithm and implementation choices.
7:45 – 8:15 PM Jason Mancuso, MTS, Research @ Modal
Elastic, Disaggregated RL Rollouts with Stitch
In fully async RL, rollout inference can require 2–4x the compute of the trainer, ballooning the minimum cluster size for a run. Disaggregating rollouts onto an elastic pool helps, but creates a new problem: getting fresh weights to a changing set of inference workers without spending a ton of time syncing or letting rollouts get too stale.
Stitch is a trainer and engine-agnostic framework for doing this. It relies on sparse delta compression to make weight sync cheap enough for rollout workers to scale independently from the training cluster. We’ll talk about why we built Stitch, its weight sync protocol, and why we’re excited about disaggregated rollouts as a better primitive for RL infra.
8:15 – 9:00 PM Discussion & networking