

Tianyi Alex Qiu | Computing Human Ideal Preference, Breaking AI Feedback Loops
Foresight Institute’s Computation Group
Computing Human Ideal Preference, Breaking AI Feedback Loops
Abstract: How should an AI system assist someone whose beliefs and preferences are still forming while under the influence of such AI systems? Classical approaches based on preference learning and inverse reinforcement learning often mistake transient/instrumental preferences (e.g., money) for stable/terminal ones (e.g., well-being). The latter is stationary while the former changes over time. By assuming the stationarity of both, classical methods are shown to entrench transient/instrumental beliefs and preferences of the human principal, which they would otherwise come to regret. To avoid this problem, we remove the stationarity assumption and propose the alternative goal of learning the human’s ideal preference, i.e., the human belief/preference state that's maximally stable upon sufficient reflection and in face of Socratic counter-persuasion. We show how this target can be formalized and computed in practice, when training language model assistants.
Bio: Tianyi co-leads Prevail, an independent group addressing epistemic disempowerment. They develop training/eval interventions (often with humans in the loop) and systemic measures, targeting failure modes like long-term value lock-in and enabling good outcomes like moral/knowledge progress. Tianyi is a researcher at Oxford HAI Lab, 0th-year PhD student at Stanford CS, and former Anthropic AI Safety Fellow. Research projects he led have received two Best Paper Awards: one at ACL'25 and one at the NeurIPS'24 Pluralistic Alignment Workshop. https://tianyiqiu.net/
A group of scientists, engineers, and entrepreneurs in computer science, ML, cryptography, and related fields who leverage those technologies to improve voluntary cooperation across humans, and ultimately AIs.
Subscribe to our newsletter
Nominate a seminar presenter/topic
Share this application with colleagues who’d like to join
Our book Gaming the Future: Technologies for Intelligent Voluntary Cooperation is now online on Substack.
Zoom link: https://us02web.zoom.us/j/81623367983