Cover Image for AI Safety Evals - Paper Reading Club
Cover Image for AI Safety Evals - Paper Reading Club
Avatar for BlueDot Impact
Presented by
BlueDot Impact
We’re building the workforce needed to safely navigate AGI.
Contact: team@bluedot.org

AI Safety Evals - Paper Reading Club

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Andreas Turanski (co-host) presents the very recent paper attempting to mitigate Evaluation Awareness to improve Evaluation Accuracy, as a follow-on to our situational-awareness series. Predicting LLM Safety Before Release by Simulating Deployment (Korbak et al., OpenAI, 2026-06-16). Blog 🔗, Paper 🔗.

  • Replays real user conversations through a new model to forecast misbehavior before release

  • Headline: evals trigger model eval-awareness ~99% vs ~5% in production; simulation deployment matches production.

Every week, someone will present for up to 20 minutes, followed by 40 minutes of discussion. RSVP to join, sign up to present, or contact us at evalsreadinggroup@gmail.com with questions. Everyone is welcome!

Avatar for BlueDot Impact
Presented by
BlueDot Impact
We’re building the workforce needed to safely navigate AGI.
Contact: team@bluedot.org