Cover Image for When an Aligned AI Still Goes Wrong
Cover Image for When an Aligned AI Still Goes Wrong
Avatar for Lorong AI
Presented by
Lorong AI
Hosted By
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

What if an AI is perfectly obedient... and that's exactly the problem?

An AI can faithfully follow every instruction from its operator and still produce outcomes that society dislikes. If that's possible, then perhaps AI alignment isn't just about building better AI systems. Perhaps it's also about the organisations, incentives and institutions behind them. This is the central idea behind full-stack alignment: aligning not only AI systems, but the wider ecosystem they operate within.

Let's read and discuss "Full-Stack Alignment: From AI Alignment to Societal Alignment", which argues that aligning AI requires looking beyond the model itself. The authors propose that today's approaches to representing human values are too simplistic, and introduce "thick models of value" as a richer way for AI to reason about values, norms and collective interests.

Then we'll take the discussion beyond the paper. If society itself rarely agrees on what is right, should AI be expected to? Who decides what an AI ought to optimise for? And when organisations and society disagree, who should AI ultimately serve?


❗Please Read Before You Register❗

This isn't a seminar or lecture. It's a discussion, and everyone in the room is part of it.

  • No expertise required. Just curiosity and a willingness to think out loud together. Whether you're AI-curious, work with AI, or simply interested in the topic, you're welcome.

  • We'll do the bulk of the reading during the session together! Then we'll use it as a shared starting point for the discussion, so coming prepared will help everyone get more out of the conversation.

  • Bring your own perspectives. If you have relevant experiences, examples from your work, or other research that challenges or adds to the paper, we'd love to hear them.


More About the Host

Jared Cheang is an AI Safety researcher with a background in data science, software engineering, and building AI solutions for the public sector. His current research interests span AI control, alignment, and safety-by-design, with a particular interest in how increasingly capable AI systems can remain reliable and aligned with human values. Jared also teaches and mentors across the local tech community, and holds a Master's in Business Analytics from NUS.


More About the Series

Paper Club is Lorong AI’s community-driven initiative where members gather to discuss and analyze academic papers, research articles, or key developments in AI.

Get involved: Learn more about Lorong AI | Speaker Sign-up | WhatsApp Community | LinkedIn | X

Location
Lorong AI @ One-North
69 Ayer Rajah Cres., Singapore 139961
The Markdown (Entrance C)
Avatar for Lorong AI
Presented by
Lorong AI
Hosted By