

Foresight's AI Salon: Georg Lange | Tracing LLMs' thoughts to detect deception, intent, and misalignment
Join Foresight Institute and Georg Lange on his session entitled Tracing LLMs' thoughts to detect deception, intent, and misalignment
Schedule
6 - 7 pm: Drinks & Social
7 - 8:30pm: Talk, Q&A, & Breakouts
8:30 - 9 pm: Mingle & Hang
Speaker Bio: "Georg is an independent researcher working on Mechanistic Interpretability for LLMs. He works on circuit tracing, dictionary learning, and automated interpretability to create technology broadly useful for detecting deception, intent, and misalignment.
Before turning to language models, Georg studied biological ones, measuring neural activations in mice to understand reward learning at the circuit level. He sees mechanistic interpretability as a rare opportunity in the science of intelligent systems: for the first time, we can record every unit, intervene on any connection, and rerun the same computation under controlled perturbations. His work uses that leverage to map what LLMs actually compute, from decomposing representations into interpretable features and tracing circuits, to detecting where a model's internal computation diverges from its stated reasoning"
Salon discussions are held under Chatham House Rule (don’t connect people to ideas when discussing them outside of this workshop).
Foresight’s AI Nodes offer grant funding, local compute and community hubs for AI for Science and Safety projects in Berlin and the Bay Area.