Cover Image for Simple and Effective RL for Recommender Systems & Diffusion Fine-tuning
Cover Image for Simple and Effective RL for Recommender Systems & Diffusion Fine-tuning
Avatar for Lossfunk Event Calendar
Your friendly neighborhood AI lab
51 Went

Simple and Effective RL for Recommender Systems & Diffusion Fine-tuning

Registration
Past Event
Welcome! Please choose your desired ticket type:
About Event

About the talk: Foundational models don’t become useful by pre-training alone; they need careful post-training and feedback to align with what users want. This talk explores how RL connects two domains: learning from logged user interactions in recommender/search systems and fine-tuning text-to-image diffusion models for better instruction following.

We’ll cover: •⁠ ⁠Why off-policy evaluation and learning are tricky in recommender systems •⁠ ⁠How closed-form baseline corrections reduce variance and improve learning •⁠ ⁠When to use REINFORCE vs PPO for post-training foundational models •⁠ ⁠LOOP: a multi-trajectory method for stable diffusion fine-tuning •⁠ ⁠Empirical results on improving attribute binding, spatial relations, and aesthetics

About the speaker: Shashank is a final-year Ph.D. student at the University of Amsterdam. His research focuses on off-policy learning from user interactions in recommender systems and RL post-training for foundational models. He has also interned at Meta AI and worked as a data scientist at Flipkart.
Linkedin:
Twitter:

Location
Indiranagar
Bengaluru, Karnataka, India
The exact location will be shred once the invite is accepted
Avatar for Lossfunk Event Calendar
Your friendly neighborhood AI lab
51 Went