Fine Tuning LLMs: Group Relative Policy Optimization
Hosted by Joel & Network School
Registration
Past Event
About Event
We walk through the process of supervised fine-tuning QWEN2.5 Math 1.5B on a grade school math dataset.
We will cover an overview of how to write a vanilla Group Relative Policy Optimization (GRPO) implementation and then demonstrate how to fine tune via the Tinker API.
We will:
1. Evaluate zero-shot Qwen2.5 Math 1.5B via vLLM and script
2. Walk through GRPO and fine-tune Qwen2.5
3. Review performance improvements
4. Demonstrate fine tuning via Tinker API
The talk assumes familiarity with basic python programming (e.g. imports, functions) and high school math (e.g. log, exp, summation)
Location
Network School Library