

One Bias After Another: Fixing Language Reward Models
To make AI systems helpful, we use reward models, which are essentially an automated grading system that tells the AI what a good answer looks like. But what happens when the grader itself is fundamentally biased?
On Wednesday, May 6th, Safe AI Germany (SAIGE) is hosting Daniel Fein & Max Lamparth from Stanford University to present their latest findings on the hidden flaws inside frontier AI models (Paper: https://arxiv.org/pdf/2603.03291)
What We Will Cover:
In this presentation and Q&A session, we will explore why state-of-the-art AI assistants remain plagued by simple biases, including
Excessively long and short responses: Why AI graders have started overcorrecting, actively punishing correct, detailed answers in favour of short, incorrect ones.
Avoiding uncertainty: Why AI prefers a confident lie over a truthful "I'm not sure."
The agreeable problem: Why would AI change its answer just to agree with the user, even when it knows the user is wrong?
Working towards solutions.
🎤 Format: 40-minute presentation followed by a 20-minute open Q&A.
Who Should Attend?
This session is designed for AI safety researchers, STEM students, and anyone interested in the technical black box of why AI makes mistakes and how to fix them. We will provide a short introduction to reward hacking before deep-diving into any technicalities.
📅 Date & Time: Wednesday, May 6th | 18:00 - 19:00 CEST
Looking forward to seeing you there!