

Beyond RLHF: MPO as the Next Generation of LLM Post-Training
RL is widely used for post-training LLMs, but we rarely question how it compares to newer contrastive and regression-based methods.
DPO made alignment simpler by removing the need for a separate reward model, but it has its own limits and biases.
Multi-Preference Optimization (MPO) takes the next step by learning from groups of preferences instead of just pairs, capturing more signal from the same data.
This leads to major performance gains on benchmarks like AlpacaEval 2, especially for open-source models.
Active MPO (AMPO) improves data efficiency by training only on the most useful examples, reducing compute cost.
Reference-Free Alignment (REFA) uses probabilistic signals to fix issues like brevity bias, without hurting response quality.
These new methods, presented at ICML, COLM, and NeurIPS 2025, show how alignment can be made more scalable, robust, and controllable.
About the speaker:
Rahul Madhavan (aka maddy), Researcher at Google Deepmind
rahul-madhavan || imrahulmaddy
Preread:
MPO (Neurips 2025 Workshop): https://arxiv.org/abs/2412.04628
AMPO (ICML 2025): https://arxiv.org/abs/2502.18293
REFA (COLM 2025): https://openreview.net/pdf?id=zP6DJaBBcR
To attend online, please use the link below:
https://meet.google.com/gck-gvax-wxa?hs=122&authuser=0