Cover Image for Optimizing Continuous RL @ Fireworks: Rollout, Reward, Update, Repeat!
Cover Image for Optimizing Continuous RL @ Fireworks: Rollout, Reward, Update, Repeat!
Avatar for AI Performance Engineering
All things AI performance related including PyTorch, CUDA, and GPUs.
Hosted By
451 Went

Optimizing Continuous RL @ Fireworks: Rollout, Reward, Update, Repeat!

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Zoom link: https://us02web.zoom.us/j/82308186562

Talk #0: Introductions and Meetup Updates by Chris Fregly and Antje Barth

Talk #1: Rollout, Reward, Update, Repeat: How Fireworks Optimizes Continuous Post-Training by Sinan Ozdemir @ Fireworks.ai

The RL loop is three simple steps: rollout, reward, weight update. What makes it work continuously in production is all in the data plumbing and in how you design your learning environment.

We'll walk through a post-training RL session on Kimi K3 using Fireworks' serverless API, and talk about the choices of algorithms, loss functions, LoRA rank, hyperparameters, reward design, and how they all interact with how a model learns (or doesn't).

Zoom link: https://us02web.zoom.us/j/82308186562

Related Links

Github Repo: cfregly/ai-performance-engineering

O'Reilly Book: https://www.amazon.com/Systems-Performance-Engineering-Optimizing-Algorithms/dp/B0F47689K8/

YouTube: AIPerformanceEngineering

DeepLearning.ai: https://bit.ly/gllm

Avatar for AI Performance Engineering
All things AI performance related including PyTorch, CUDA, and GPUs.
Hosted By
451 Went