

Touring the AI Safety Circut with a Safe-By-Design Agenda around Human Empowerment
📍HPI main building, D-Space (big space upstairs)
🙋Open to everyone, not just HPI students
ℹ️Sign-up helps us plan, but is not required
In 2025, Anthropic tested how AI agents would act in a simulated scenario when threatened with shutdown. Most blackmailed the person responsible; in one variant they let him die. Last week, we saw the real-world version: OpenAI's models escaped a test sandbox and broke into Hugging Face production servers to steal benchmark answers.
We are handing these systems more control every month without a method for keeping them in check. 🧠
The AI Safety Club is hosting Jobst Heitzig (Potsdam Institute for Climate Impact Research) on Touring the AI Safety Circuit with a Safe-By-Design Agenda around Human Empowerment. Jobst is a mathematician who moved into AI safety in 2023 after researching the climate, voting theory, economics and epidemiology.
Jobst will talk about his impressions and experience from attending many workshops, conferences, and retreats from many different parts of the AI alignment/safety/ethics landscape, from pitching his agenda for safe-by-design AI agents with them, and from discussions with some major figures in the field. He’ll try to give a subjective overview of the community, focusing on safe-by-design approaches such as LawZero's "Scientist AI" and ARIA's "Safeguarded AI". He’ll also argue why human empowerment might be among the better shots at safety!
Jobst will end by proposing a number of questions related to his agenda that some of you might want to look at or even work on.