When an Aligned AI Still Goes Wrong
What if an AI is perfectly obedient... and that's exactly the problem?
An AI can faithfully follow every instruction from its operator and still produce outcomes that society dislikes. If that's possible, then perhaps AI alignment isn't just about building better AI systems. Perhaps it's also about the organisations, incentives and institutions behind them. This is the central idea behind full-stack alignment: aligning not only AI systems, but the wider ecosystem they operate within.
Let's read and discuss "Full-Stack Alignment: From AI Alignment to Societal Alignment", which argues that aligning AI requires looking beyond the model itself. The authors propose that today's approaches to representing human values are too simplistic, and introduce "thick models of value" as a richer way for AI to reason about values, norms and collective interests.
Then we'll take the discussion beyond the paper. If society itself rarely agrees on what is right, should AI be expected to? Who decides what an AI ought to optimise for? And when organisations and society disagree, who should AI ultimately serve?
❗Please Read Before You Register❗
This isn't a seminar or lecture. It's a discussion, and everyone in the room is part of it.
No expertise required. Just curiosity and a willingness to think out loud together. Whether you're AI-curious, work with AI, or simply interested in the topic, you're welcome.
We'll do the bulk of the reading during the session together! Then we'll use it as a shared starting point for the discussion, so coming prepared will help everyone get more out of the conversation.
Bring your own perspectives. If you have relevant experiences, examples from your work, or other research that challenges or adds to the paper, we'd love to hear them.
More About the Host
Jared Cheang is an AI Safety researcher with a background in data science, software engineering, and building AI solutions for the public sector. His current research interests span AI control, alignment, and safety-by-design, with a particular interest in how increasingly capable AI systems can remain reliable and aligned with human values. Jared also teaches and mentors across the local tech community, and holds a Master's in Business Analytics from NUS.
More About the Series
Paper Club is Lorong AI’s community-driven initiative where members gather to discuss and analyze academic papers, research articles, or key developments in AI.
Get involved: Learn more about Lorong AI | Speaker Sign-up | WhatsApp Community | LinkedIn | X
