Cover Image for Beyond RLHF: MPO as the Next Generation of LLM Post-Training
Cover Image for Beyond RLHF: MPO as the Next Generation of LLM Post-Training
Avatar for Lossfunk Event Calendar
Your friendly neighborhood AI lab
46 Went

Beyond RLHF: MPO as the Next Generation of LLM Post-Training

Register to See Address
Bengaluru, India
Registration
Past Event
Welcome! Please choose your desired ticket type:
About Event
  • RL is widely used for post-training LLMs, but we rarely question how it compares to newer contrastive and regression-based methods.

  • DPO made alignment simpler by removing the need for a separate reward model, but it has its own limits and biases.

  • Multi-Preference Optimization (MPO) takes the next step by learning from groups of preferences instead of just pairs, capturing more signal from the same data.

  • This leads to major performance gains on benchmarks like AlpacaEval 2, especially for open-source models.

  • Active MPO (AMPO) improves data efficiency by training only on the most useful examples, reducing compute cost.

  • Reference-Free Alignment (REFA) uses probabilistic signals to fix issues like brevity bias, without hurting response quality.

  • These new methods, presented at ICML, COLM, and NeurIPS 2025, show how alignment can be made more scalable, robust, and controllable.

About the speaker:
Rahul Madhavan (aka maddy), Researcher at Google Deepmind
||

Preread:
MPO (Neurips 2025 Workshop): https://arxiv.org/abs/2412.04628
AMPO (ICML 2025): https://arxiv.org/abs/2502.18293
REFA (COLM 2025): https://openreview.net/pdf?id=zP6DJaBBcR

To attend online, please use the link below:
https://meet.google.com/gck-gvax-wxa?hs=122&authuser=0

Location
Please register to see the exact location of this event.
Bengaluru, India
Avatar for Lossfunk Event Calendar
Your friendly neighborhood AI lab
46 Went