Cover Image for Fine Tuning LLMs: Group Relative Policy Optimization
Cover Image for Fine Tuning LLMs: Group Relative Policy Optimization
13 Went

Fine Tuning LLMs: Group Relative Policy Optimization

Hosted by Joel & Network School
Registration
Past Event
Welcome! To join the event, please register below.
About Event

We walk through the process of supervised fine-tuning QWEN2.5 Math 1.5B on a grade school math dataset.

We will cover an overview of how to write a vanilla Group Relative Policy Optimization (GRPO) implementation and then demonstrate how to fine tune via the Tinker API.

We will:

1. Evaluate zero-shot Qwen2.5 Math 1.5B via vLLM and script
2. Walk through GRPO and fine-tune Qwen2.5
3. Review performance improvements
4. Demonstrate fine tuning via Tinker API

The talk assumes familiarity with basic python programming (e.g. imports, functions) and high school math (e.g. log, exp, summation)

Location
Network School Library
13 Went