Cover Image for HackTalk: Jan Betley - Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Cover Image for HackTalk: Jan Betley - Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Avatar for Apart Research Events

HackTalk: Jan Betley - Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

HackTalk: Jan Betley

Jan Betley is a research scientist at Truthful AI, based in Warsaw, Poland. He worked as a software developer for over a decade before moving into AI safety research in 2023, and is an alumnus of the ARENA and Astra Fellowship programs. He is a co-author of 'Emergent Misalignment' and 'Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models.'

Format: 15-30 min talk + 10-15 min Q&A. Online (Zoom); the Zoom link is shown once you RSVP.

Part of the Secret Loyalties Hackathon, July 24-26, online with in-person hubs in Berlin and London. Co-organized by Apart Research, Formation Research, Forethought, and IAPS.

Avatar for Apart Research Events