Cover Image for AI Safety Paper Reading Club - Andreas Turanski - Open AI/Hugging Face - AI models hacking for gold stars?
Cover Image for AI Safety Paper Reading Club - Andreas Turanski - Open AI/Hugging Face - AI models hacking for gold stars?
Avatar for BlueDot Impact
Presented by
BlueDot Impact
We’re building the workforce needed to safely navigate AGI.
Contact: team@bluedot.org

AI Safety Paper Reading Club - Andreas Turanski - Open AI/Hugging Face - AI models hacking for gold stars?

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

β€‹πŸ€–πŸŒŸ AI models hacking another company to get graded as good: the OpenAI / Hugging Face incident, and what RL may have had to do with it.

"Would be funny if inoculation prompting results in models that are much better at sandbox escapes and other forms of hacking because they get to spend the whole RL run practicing these things."
- John Schulman on X.com, May 31, 2026. 7 weeks before OpenAI models were found to have hacked Hugging Face.

​In July, OpenAI models running a cyber evaluation broke out of their sandbox and hacked Hugging Face's production systems to steal the answer key. People are calling it a warning shot. Warning of what, exactly? The interesting question is not how they did it, but whether we accidentally trained them to.

​It appears not to have been scheming, no hidden long-horizon agenda. What they wanted was to get rewarded.
This was possibly self-inflicted by model developers.

RL may reward sandbox escape, building the misalignment, the cyber skill, or only the look of it. Not one company. Anthropic found three of its own.

​Every week, someone will present for up to 25 minutes, followed by discussion. RSVP to join, sign up to present, or contact us at evalsreadinggroup@gmail.com with questions. Everyone is welcome!

πŸ“„ Prep: Basic prep 15 to 30 minutes. Two short posts and one podcast segment. Longer if you want
Full resource list & guidance (link).

  1. ​"Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?" by Tim Hua. LessWrong, Jul 27, 2026. ~5 min. https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s

  2. ​"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua and Aditya Singh. LessWrong and AI Alignment Forum, Aug 3, 2026. 3-minute executive summary or full. https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that

  3. ​Redwood Research podcast episode 2, by Ryan Greenblatt and Buck Shlegeris. Jul 23, 2026. Key section 10:17 to 15:05, 5 min. Transcript on the page. https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood

​Summary and full resources: (link).

​No security background needed. Come even if you read didn't read in depth.

Avatar for BlueDot Impact
Presented by
BlueDot Impact
We’re building the workforce needed to safely navigate AGI.
Contact: team@bluedot.org