CAIA Reading Group: Concrete Problems in AI Safety, & the OpenAI / HuggingFace Incident
Join us for the first CAIA reading group discussion of the semester!
This week’s discussion accompanies Week 1 of CS 1998: Introduction to AI Safety & Alignment. We’ll discuss concrete problems in AI safety—how failures arise in real AI systems, how the field has historically framed these problems, and what they look like in today’s increasingly agentic systems.
We’ll start with the classic 2016 paper “Concrete Problems in AI Safety” by Dario Amodei et al., which introduces problems including reward hacking, scalable oversight, safe exploration, robustness to distributional shift, and avoiding negative side effects.
We’ll then look at a recent real-world case study: the OpenAI / Hugging Face incident, using reports from OpenAI and METR to discuss how concrete safety problems can manifest in modern AI agents interacting with real systems.
Monday, August 31 · 4:30–5:30 PM
Led by: Arya Datla
Location: CIS Building 350
🍽️ Food will be provided!
Readings:
Dario Amodei et al., Concrete Problems in AI Safety https://arxiv.org/abs/1606.06565
OpenAI, Hugging Face Incident and the Road Ahead https://openai.com/index/hugging-face-incident-and-the-road-ahead/
METR, OpenAI–Hugging Face Incident Investigation https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation
You don’t need to be enrolled in CS 1998 to attend, all are welcome!