

90/30 Club (ML reading) #52: DeepSeek-V4: Million Token Context and the Next Frontier of Test-Time Scaling
Week 52: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4 represents a shift from scaling model parameters to scaling context and computation efficiency. By introducing hybrid attention (Compressed Sparse Attention + Heavily Compressed Attention), manifold-constrained residual pathways, and a new optimizer (Muon), the paper shows how LLMs can efficiently operate over million-token contexts, unlocking long-horizon reasoning, agent workflows, and test-time scaling.
Rather than being bottlenecked by quadratic attention, DeepSeek reframes progress as an efficiency problem: compress memory (KV cache), reduce FLOPs, and enable sustained reasoning over massive sequences. This allows models to handle tasks like multi-document synthesis, long-running agents, and real-world workflows that were previously infeasible.
Join us at Mox to explore:
• Does scaling context length (vs parameters) represent the next dominant axis of LLM progress?
• How do compression-based attention mechanisms (CSA/HCA) change the tradeoff between memory, latency, and reasoning depth?
🔎Analyzed Papers
✍Google Drive for sharing Comments
Discussion at 20:00, (optional) quiet reading from 19:00.