Cover Image for AI Safety Technical Reading Group – Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Cover Image for AI Safety Technical Reading Group – Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Communauté montréalaise pour la sécurité, l’éthique et la gouvernance de l’IA. // Montréal community for AI safety, ethics, and governance.

AI Safety Technical Reading Group – Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Get Tickets
Welcome! Please choose your desired ticket type:
About Event

​AI Safety Technical Reading Group

​Each month, we gather to discuss an important technical AI safety paper. We recommend reading the paper before and to come with your commentary and insights. Discussion bilingue.

​Groupe de lecture sur la sécurité technique de l'IA

​Chaque mois, nous nous rencontrons pour discuter un papier important en sécurité technique de l'IA. Nous recommandons lire le papier avant et de venir avec votre commentaire et bonnes idées.

​The paper this time:

​Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values (Betley et al., Truthful AI, 2026) — https://arxiv.org/abs/2607.14345

​People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user.

​In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user.

​Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others.

​We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.

​Où / Where

​AI Safety Montréal and Ω Labs run on member contributions. Your donation keeps the project going and the snacks coming.

Location
Omega Labs
3813 R. Saint-Denis, Montréal, QC H2W 2M4, Canada
Hybrid event, also join us online via Zoom (at 6:15 PM): https://zoom.us/j/99005830598?pwd=3e21GRcTVAPTbvGO…
Communauté montréalaise pour la sécurité, l’éthique et la gouvernance de l’IA. // Montréal community for AI safety, ethics, and governance.