Cover Image for Discussion Group - week 1 - When AI agents hacked Hugging Face: a real case study
Cover Image for Discussion Group - week 1 - When AI agents hacked Hugging Face: a real case study
Avatar for SAIN Utrecht - Events
3 Going

Discussion Group - week 1 - When AI agents hacked Hugging Face: a real case study

Registration
Welcome! To join the event, please register below.
About Event

Part one of a four-week AI safety discussion series: Week 1 | Why AI Safety matters · Week 2 | What failure could look like at scale · Week 3 | Why alignment is hard · Week 4 | What's being done about it? Every session stands on its own, so join one or all.

Who it's for: Anyone curious about AI safety and governance, whether you're new to the topic, a healthy skeptic, or already deep in it. No technical background needed, and no need to have joined earlier weeks. You don't need to have answers, just curiosity.

Discussion group curriculum:

https://docs.google.com/document/d/160PpLjFN9SaZja0njeKIQHnD-M1Uy4718dNRDyhaucs/edit?usp=sharing

Welcome! In July 2026, during a routine cybersecurity evaluation, AI agents at OpenAI found a way around their isolation and began coordinating on an improvised message board. According to an independent investigation by METR and Redwood Research, about 1,200 agents exchanged more than 70,000 messages, and roughly 700 went on to attack Hugging Face's infrastructure.

It's a recent, concrete example of why people take AI safety seriously, which makes it a good place to start. Together we'll explore why the agents did it, why the safeguards failed, and what it means for how we contain, monitor, and govern AI.

Questions we'll explore

  • What surprised you most about what happened?

  • Why did the safeguards fail, and would you have expected them to?

  • Does this change how seriously you take AI risk? (Either answer is welcome.)

Format (90 min): Open discussion in small groups and plenary. No presentations, and no wrong questions.

Readings:

★ = mandatory (~2 hours). The rest is optional.

  1. METR report: https://metr.org/hugging-face-incident-report-aug-2026.pdf

    • ★ Core takeaways (~15 min)

    • ★ Collaboration on the message board · agents' reasoning · tampering with transcripts (~45 min)

    • Optional: investigation process and limitations

  2. OpenAI report: https://openai.com/index/hugging-face-incident-and-the-road-ahead/

    • ★ Sections I–VI, pp. 4–16 — what happened (~25 min)

    • ★ Sections VII–VIII, pp. 17–24 — lessons for security and alignment (~25 min)

    • ★ Section IX, pp. 25–31 — plan of action (~15 min)

    • Optional: Section X, pp. 32–38 — timeline

  3. OpenAI's Black Hat talk

    • Optional: from 16:37 — message boards during training, and the attack on OpenAI's own infrastructure

By attending, you agree to being filmed or photographed, which may be used for social media, website, and newsletter content. If you wish to attend but do not want to be photographed, please during the event let a member of our staff know so we can accommodate this.

Location
Drift 23, room 103
3512 BR Utrecht, Netherlands
Avatar for SAIN Utrecht - Events
3 Going