

HackTalk: Jan Betley - Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Registration
Past Event
About Event
HackTalk: Jan Betley
Jan Betley is a research scientist at Truthful AI, based in Warsaw, Poland. He worked as a software developer for over a decade before moving into AI safety research in 2023, and is an alumnus of the ARENA and Astra Fellowship programs. He is a co-author of 'Emergent Misalignment' and 'Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models.'
Format: 15-30 min talk + 10-15 min Q&A. Online (Zoom); the Zoom link is shown once you RSVP.
Part of the Secret Loyalties Hackathon, July 24-26, online with in-person hubs in Berlin and London. Co-organized by Apart Research, Formation Research, Forethought, and IAPS.