Cover Image for CAIA Reading Group: Concrete Problems in AI Safety, & the OpenAI / HuggingFace Incident
Cover Image for CAIA Reading Group: Concrete Problems in AI Safety, & the OpenAI / HuggingFace Incident
Avatar for Cornell AI Alignment
A community of students and researchers conducting research and outreach to mitigate risks from advanced AI systems.
Hosted By
24 Went

CAIA Reading Group: Concrete Problems in AI Safety, & the OpenAI / HuggingFace Incident

Registration
Past Event
Welcome! To join the event, please register below.
About Event

Join us for the first CAIA reading group discussion of the semester!

This week’s discussion accompanies Week 1 of CS 1998: Introduction to AI Safety & Alignment. We’ll discuss concrete problems in AI safety—how failures arise in real AI systems, how the field has historically framed these problems, and what they look like in today’s increasingly agentic systems.

We’ll start with the classic 2016 paper “Concrete Problems in AI Safety” by Dario Amodei et al., which introduces problems including reward hacking, scalable oversight, safe exploration, robustness to distributional shift, and avoiding negative side effects.

We’ll then look at a recent real-world case study: the OpenAI / Hugging Face incident, using reports from OpenAI and METR to discuss how concrete safety problems can manifest in modern AI agents interacting with real systems.

Monday, August 31 · 4:30–5:30 PM
Led by: Arya Datla
Location: CIS Building 350
🍽️ Food will be provided!

Readings:
Dario Amodei et al., Concrete Problems in AI Safety https://arxiv.org/abs/1606.06565
OpenAI, Hugging Face Incident and the Road Ahead https://openai.com/index/hugging-face-incident-and-the-road-ahead/
METR, OpenAI–Hugging Face Incident Investigation https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation

You don’t need to be enrolled in CS 1998 to attend, all are welcome!

Location
CIS Building 350
Avatar for Cornell AI Alignment
A community of students and researchers conducting research and outreach to mitigate risks from advanced AI systems.
Hosted By
24 Went