

AI Safety Paper Reading Club - Andreas Turanski - Open AI/Hugging Face - AI models hacking for gold stars?
βπ€π AI models hacking another company to get graded as good: the OpenAI / Hugging Face incident, and what RL may have had to do with it.
"Would be funny if inoculation prompting results in models that are much better at sandbox escapes and other forms of hacking because they get to spend the whole RL run practicing these things."
- John Schulman on X.com, May 31, 2026. 7 weeks before OpenAI models were found to have hacked Hugging Face.
βIn July, OpenAI models running a cyber evaluation broke out of their sandbox and hacked Hugging Face's production systems to steal the answer key. People are calling it a warning shot. Warning of what, exactly? The interesting question is not how they did it, but whether we accidentally trained them to.
βIt appears not to have been scheming, no hidden long-horizon agenda. What they wanted was to get rewarded.
This was possibly self-inflicted by model developers.
RL may reward sandbox escape, building the misalignment, the cyber skill, or only the look of it. Not one company. Anthropic found three of its own.
βEvery week, someone will present for up to 25 minutes, followed by discussion. RSVP to join, sign up to present, or contact us at evalsreadinggroup@gmail.com with questions. Everyone is welcome!
π Prep: Basic prep 15 to 30 minutes. Two short posts and one podcast segment. Longer if you want
Full resource list & guidance (link).
β"Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?" by Tim Hua. LessWrong, Jul 27, 2026. ~5 min. https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s
β"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua and Aditya Singh. LessWrong and AI Alignment Forum, Aug 3, 2026. 3-minute executive summary or full. https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that
βRedwood Research podcast episode 2, by Ryan Greenblatt and Buck Shlegeris. Jul 23, 2026. Key section 10:17 to 15:05, 5 min. Transcript on the page. https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood
βSummary and full resources: (link).
βNo security background needed. Come even if you read didn't read in depth.