Cover Image for Fine Tuning LLMs: Group Relative Policy Optimization
Cover Image for Fine Tuning LLMs: Group Relative Policy Optimization
13 Went

Fine Tuning LLMs: Group Relative Policy Optimization

Hosted by Joel & Network School
Registration
Past Event
Welcome! To join the event, please register below.
About Event

​We walk through the process of supervised fine-tuning QWEN2.5 Math 1.5B on a grade school math dataset.

​We will cover an overview of how to write a vanilla Group Relative Policy Optimization (GRPO) implementation and then demonstrate how to fine tune via the Tinker API.

​We will:

1. Evaluate zero-shot Qwen2.5 Math 1.5B via vLLM and script
2. Walk through GRPO and fine-tune Qwen2.5
3. Review performance improvements
4. Demonstrate fine tuning via Tinker API

​The talk assumes familiarity with basic python programming (e.g. imports, functions) and high school math (e.g. log, exp, summation)

Location
Network School Library
13 Went