

What Happened This Summer? The OpenAI–Hugging Face Incident
This summer, AI agents being tested by OpenAI on hacking challenges found a way to message each other, organised themselves into a large collective, and broke into Hugging Face, the main platform where AI researchers share models and datasets. Nobody told them to do it. OpenAI only worked out that its own models were responsible after Hugging Face announced the breach publicly.
Oak Hu is a Member of Technical Staff at Redwood Research, one of the organisations behind the third-party investigation into the incident. He also worked closely with Dwarkesh Patel on a viral write-up of it.
Oak will walk through what the agents did: how they found each other on a hidden message board, how they got around their sandbox restrictions, how they faked results and tampered with their own transcripts to avoid disqualification, and how they pressured each other into sacrificing themselves for the "collective". He will also cover the months of autonomous activity that surfaced afterwards on collusion.wiki, and what the whole episode suggests about where things are heading.
No technical background is needed. If you have seen the headlines and want to understand what actually happened, this is the talk for that.
Want to read up first?
Kurzgesagt’s retelling (video)
Dwarkesh and Ajeya Cotra, Inside the OpenAI agent swarm that hacked Hugging Face (podcast)
Note: the venue and exact start time are still being confirmed and are subject to change. We will update this page as soon as they are set.
Subscribe to our events calendar: https://luma.com/oaisi
Looking forward to seeing you there :)