Papers We Love: Brooklyn - DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Papers We Love and Espresso AI are pleased to present the second of our summer series. Papers We Love is a community of programmers who love reading and discussing computer science papers.
What was the last paper within the realm of computing you read and loved? What did it inspire you to build or tinker with? Come share the ideas in an awesome academic or research paper with fellow engineers, programmers, and paper-readers. Lead a session and show off code you wrote that implements these ideas, or just give us the lowdown on the paper. Otherwise, just come, listen, learn, and discuss.
Speaker bio:
Edwin has spent 6+ years as an ML engineer building large-scale search engines, recommender systems, and natural language models at Zocdoc and Goldbelly. At Espresso AI, he's building models of how workloads collide and contend with each other on shared compute — and cutting people's cloud bills in half in the process. When he's not yapping at Claude, he's singing (R&B, jazz, and Indian classical music) and producing a YouTube series where LLMs play poker and trash-talk each other.
Abstract:
Reward the Answer, Not the Work: DeepSeek-R1 18 Months Later
In 2023, frontier model companies stopped sharing their secret sauce. GPT-4's technical report disclosed no architecture, no compute, no training method. o1 arrived a year later having found a way to make models reason -- not only was there no explanation, but they even hid the chain of thought. Four months after, DeepSeek, a lab spun out of a Chinese quant fund, finally published the recipe and showed how models can learn to reason.
Their claim was subtractive: no process reward model, no Monte Carlo tree search, no supervised warm start. Ask a base model math questions, give it 1 if the answer is right and 0 otherwise, and it starts thinking for longer and checking its own work.
It was legible enough to argue with and open-weights for people to build on, and people did both. We'll read R1 closely, then use it as a door into the 18 months it set off.
References
The Paper(s):
Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025). Preprint: arXiv:2501.12948
Shao, Z. et al. DeepSeekMath. arXiv:2402.03300 — where GRPO actually comes from
Responses:
Luo, M. et al. DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL. Agentica / Berkeley Sky Computing + BAIR, 2025
Muennighoff, N. et al. s1: Simple test-time scaling. arXiv:2501.19393 — 1,000 samples and budget forcing get you how far?
Liu, Z. et al. Understanding R1-Zero-Like Training: A Critical Perspective. arXiv:2503.20783
Yue, Y. et al. Does RL Really Incentivize Reasoning Capacity Beyond the Base Model? arXiv:2504.13837
Spurious Rewards: Rethinking Training Signals in RLVR. arXiv:2506.10947
Chen, Y. et al. Reasoning Models Don't Always Say What They Think. arXiv:2505.05410
Liu, M. et al. ProRL. arXiv:2505.24864
Team Kimi. Kimi k1.5. arXiv:2501.12599 — same week, different recipe, same conclusion
Khatri, D. et al. The Art of Scaling Reinforcement Learning Compute for LLMs. arXiv:2510.13786
6:30pm - Doors
7:00pm - Speakers begin
PwL curate a repository of papers and places to find them. Contributions welcome via PR.
Join the PWL Discord and hop into the #nyc channel: Join Discord Server
Papers We Love has a Code of Conduct. Be good to each other and to the PWL community.
