Papers We Love: Brooklyn - DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Hosted by Marthe Naudts & 5 others
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Papers We Love and Espresso AI are pleased to present the second of our summer series. Papers We Love is a community of programmers who love reading and discussing computer science papers.

What was the last paper within the realm of computing you read and loved? What did it inspire you to build or tinker with? Come share the ideas in an awesome academic or research paper with fellow engineers, programmers, and paper-readers. Lead a session and show off code you wrote that implements these ideas, or just give us the lowdown on the paper. Otherwise, just come, listen, learn, and discuss.

Speaker bio:

Edwin has spent 6+ years as an ML engineer building large-scale search engines, recommender systems, and natural language models at Zocdoc and Goldbelly. At Espresso AI, he's building models of how workloads collide and contend with each other on shared compute — and cutting people's cloud bills in half in the process. When he's not yapping at Claude, he's singing (R&B, jazz, and Indian classical music) and producing a YouTube series where LLMs play poker and trash-talk each other.

Abstract:

Reward the Answer, Not the Work: DeepSeek-R1 18 Months Later

In 2023, frontier model companies stopped sharing their secret sauce. GPT-4's technical report disclosed no architecture, no compute, no training method. o1 arrived a year later having found a way to make models reason -- not only was there no explanation, but they even hid the chain of thought. Four months after, DeepSeek, a lab spun out of a Chinese quant fund, finally published the recipe and showed how models can learn to reason.

Their claim was subtractive: no process reward model, no Monte Carlo tree search, no supervised warm start. Ask a base model math questions, give it 1 if the answer is right and 0 otherwise, and it starts thinking for longer and checking its own work.

It was legible enough to argue with and open-weights for people to build on, and people did both. We'll read R1 closely, then use it as a door into the 18 months it set off.

References

The Paper(s):

  • Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025). Preprint: arXiv:2501.12948

  • Shao, Z. et al. DeepSeekMath. arXiv:2402.03300 — where GRPO actually comes from

Responses:

  • Luo, M. et al. DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL. Agentica / Berkeley Sky Computing + BAIR, 2025

  • Muennighoff, N. et al. s1: Simple test-time scaling. arXiv:2501.19393 — 1,000 samples and budget forcing get you how far?

  • Liu, Z. et al. Understanding R1-Zero-Like Training: A Critical Perspective. arXiv:2503.20783

  • Yue, Y. et al. Does RL Really Incentivize Reasoning Capacity Beyond the Base Model? arXiv:2504.13837

  • Spurious Rewards: Rethinking Training Signals in RLVR. arXiv:2506.10947

  • Chen, Y. et al. Reasoning Models Don't Always Say What They Think. arXiv:2505.05410

  • Liu, M. et al. ProRL. arXiv:2505.24864

  • Team Kimi. Kimi k1.5. arXiv:2501.12599 — same week, different recipe, same conclusion

  • Khatri, D. et al. The Art of Scaling Reinforcement Learning Compute for LLMs. arXiv:2510.13786

6:30pm - Doors

7:00pm - Speakers begin

PwL curate a repository of papers and places to find them. Contributions welcome via PR.

Join the PWL Discord and hop into the #nyc channel: Join Discord Server

Papers We Love has a Code of Conduct. Be good to each other and to the PWL community.

Location
Mindspace Williamsburg
25 Kent Ave lobby; Suite 401, Brooklyn, NY 11249, USA
Use the North Lobby and go to the 4th floor.