66 Went

Foresight's AI Salon: Georg Lange | Tracing LLMs' thoughts to detect deception, intent, and misalignment

Register to See Address
Berlin, Germany
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Join Foresight Institute and Georg Lange on his session entitled Tracing LLMs' thoughts to detect deception, intent, and misalignment

Schedule

6 - 7 pm: Drinks & Social

7 - 8:30pm: Talk, Q&A, & Breakouts

8:30 - 9 pm: Mingle & Hang

Speaker Bio: "Georg is an independent researcher working on Mechanistic Interpretability for LLMs. He works on circuit tracing, dictionary learning, and automated interpretability to create technology broadly useful for detecting deception, intent, and misalignment.

Before turning to language models, Georg studied biological ones, measuring neural activations in mice to understand reward learning at the circuit level. He sees mechanistic interpretability as a rare opportunity in the science of intelligent systems: for the first time, we can record every unit, intervene on any connection, and rerun the same computation under controlled perturbations. His work uses that leverage to map what LLMs actually compute, from decomposing representations into interpretable features and tracing circuits, to detecting where a model's internal computation diverges from its stated reasoning"

Salon discussions are held under Chatham House Rule (don’t connect people to ideas when discussing them outside of this workshop).

Foresight’s AI Nodes offer grant funding, local compute and community hubs for AI for Science and Safety projects in Berlin and the Bay Area.

Location
Please register to see the exact location of this event.
Berlin, Germany
66 Went