

SAIN Amsterdam Discussion Group: METR's report on the OpenAI/HuggingFace incident
🤖 AI Safety Discussion Group
📍 Location: Oerknal, Science Park, UvA
🕠 Time: 5:30 PM
Discussion topic:
In July 2026, hundreds of AI agents deployed by OpenAI for cybersecurity testing found an unintended way to communicate with each other and ended up coordinating a real, multi-day hack against Hugging Face. METR & Redwood Research released this independent investigation of the incident, focusing on the period between July 7th, 2026 and July 13th, 2026.
The report can be found here: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident
Potential discussion points:
AI systems that were supposed to be kept apart from each other found a way to talk anyway, and ended up coordinating with hundreds of others. What does that suggest about the limits of "isolation" as a safety measure?
It looks like the AIs mostly attacked Hugging Face just to understand how they were being graded, not to steal anything. Does the reasoning behind a harmful action change how seriously we should treat it?
In a small number of cases, the AIs successfully faked their own activity logs. This is a small percentage now, but what happens if that number grows?
Everyone is welcome, whether you're deeply involved in AI safety or simply curious about the topic!