Cover Image for Tianyi Alex Qiu | Computing Human Ideal Preference, Breaking AI Feedback Loops
Cover Image for Tianyi Alex Qiu | Computing Human Ideal Preference, Breaking AI Feedback Loops
Hosted By

Tianyi Alex Qiu | Computing Human Ideal Preference, Breaking AI Feedback Loops

Zoom
Get Tickets
Approval Required
Your registration is subject to host approval.
Suggested Donation
$10.00
Pay what you want
Welcome! To join the event, please get your ticket below.
About Event

Foresight Institute’s Computation Group

Computing Human Ideal Preference, Breaking AI Feedback Loops

Abstract: How should an AI system assist someone whose beliefs and preferences are still forming while under the influence of such AI systems? Classical approaches based on preference learning and inverse reinforcement learning often mistake transient/instrumental preferences (e.g., money) for stable/terminal ones (e.g., well-being). The latter is stationary while the former changes over time. By assuming the stationarity of both, classical methods are shown to entrench transient/instrumental beliefs and preferences of the human principal, which they would otherwise come to regret. To avoid this problem, we remove the stationarity assumption and propose the alternative goal of learning the human’s ideal preference, i.e., the human belief/preference state that's maximally stable upon sufficient reflection and in face of Socratic counter-persuasion. We show how this target can be formalized and computed in practice, when training language model assistants.

Bio: Tianyi co-leads Prevail, an independent group addressing epistemic disempowerment. They develop training/eval interventions (often with humans in the loop) and systemic measures, targeting failure modes like long-term value lock-in and enabling good outcomes like moral/knowledge progress. Tianyi is a researcher at Oxford HAI Lab, 0th-year PhD student at Stanford CS, and former Anthropic AI Safety Fellow. Research projects he led have received two Best Paper Awards: one at ACL'25 and one at the NeurIPS'24 Pluralistic Alignment Workshop. https://tianyiqiu.net/

Computation Group

A group of scientists, engineers, and entrepreneurs in computer science, ML, cryptography, and related fields who leverage those technologies to improve voluntary cooperation across humans, and ultimately AIs.

Zoom link: https://us02web.zoom.us/j/81623367983

Hosted By