Cover Image for One Bias After Another: Fixing Language Reward Models
Cover Image for One Bias After Another: Fixing Language Reward Models
Avatar for SAIGE Calendar
Presented by
SAIGE Calendar
Open to everyone worldwide, unless otherwise specified.
Hosted By
148 Went

One Bias After Another: Fixing Language Reward Models

Google Meet
Registration
Past Event
Welcome! To join the event, please register below.
About Event

To make AI systems helpful, we use reward models, which are essentially an automated grading system that tells the AI what a good answer looks like. But what happens when the grader itself is fundamentally biased?

On Wednesday, May 6th, Safe AI Germany (SAIGE) is hosting Daniel Fein & Max Lamparth from Stanford University to present their latest findings on the hidden flaws inside frontier AI models (Paper: https://arxiv.org/pdf/2603.03291)

What We Will Cover:
In this presentation and Q&A session, we will explore why state-of-the-art AI assistants remain plagued by simple biases, including

  • Excessively long and short responses: Why AI graders have started overcorrecting, actively punishing correct, detailed answers in favour of short, incorrect ones.

  • Avoiding uncertainty: Why AI prefers a confident lie over a truthful "I'm not sure."

  • The agreeable problem: Why would AI change its answer just to agree with the user, even when it knows the user is wrong?

  • Working towards solutions.

🎤 Format: 40-minute presentation followed by a 20-minute open Q&A.

Who Should Attend?
This session is designed for AI safety researchers, STEM students, and anyone interested in the technical black box of why AI makes mistakes and how to fix them. We will provide a short introduction to reward hacking before deep-diving into any technicalities.

📅 Date & Time: Wednesday, May 6th | 18:00 - 19:00 CEST

Looking forward to seeing you there!

Avatar for SAIGE Calendar
Presented by
SAIGE Calendar
Open to everyone worldwide, unless otherwise specified.
Hosted By
148 Went