

90/30 Club Reading: A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
Come join us for a group discussion of "A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators"
Paper link: https://arxiv.org/abs/2604.08044
This reading group will explore ATLAS, a silicon-validated framework for evaluating LLM accelerators that stack compute directly beneath 3D DRAM. The paper shows that faster decoding requires more than maximizing memory bandwidth: memory layout, compute balance, on-chip communication, software tiling, and thermal limits must be co-designed. Its optimized cloud architecture achieves 2.53× higher decoding performance and 6.66× better energy efficiency than an NVIDIA H200 on the evaluated workloads. This is important because modern LLM generation is increasingly constrained by moving model weights and KV-cache data and 3D-integrated memory could substantially reshape inference hardware.
Event Schedule:
7pm-8pm: Quiet reading time, grab a snack and read! (optional)
8pm-9pm: Group discussion about the paper 📝
9pm-10pm: We have our space for a bit longer, stay to socialize or network!
Our event is hosted within Mox SF, the gracious donors of the space where we will meet.