

Simple and Effective RL for Recommender Systems & Diffusion Fine-tuning
About the talk: Foundational models don’t become useful by pre-training alone; they need careful post-training and feedback to align with what users want. This talk explores how RL connects two domains: learning from logged user interactions in recommender/search systems and fine-tuning text-to-image diffusion models for better instruction following.
We’ll cover: • Why off-policy evaluation and learning are tricky in recommender systems • How closed-form baseline corrections reduce variance and improve learning • When to use REINFORCE vs PPO for post-training foundational models • LOOP: a multi-trajectory method for stable diffusion fine-tuning • Empirical results on improving attribute binding, spatial relations, and aesthetics
About the speaker:
Shashank is a final-year Ph.D. student at the University of Amsterdam. His research focuses on off-policy learning from user interactions in recommender systems and RL post-training for foundational models. He has also interned at Meta AI and worked as a data scientist at Flipkart.
Linkedin: shashank-gupta-7038a223
Twitter: shashank27392