Cover Image for Paper Club: "Chain-of-Thought Is Not Explainability"
Cover Image for Paper Club: "Chain-of-Thought Is Not Explainability"
17 Went

Paper Club: "Chain-of-Thought Is Not Explainability"

Registration
Past Event
Welcome! To join the event, please register below.
About Event

This is part one of a two-part series examining the role of Chain-of-Thought (CoT) reasoning in AI safety. This week, we'll explore the current limitations of CoT as an interpretability technique, while in the next session we'll discuss how CoT can still be leveraged effectively for safety applications despite these constraints.

Technical Note: This event is intended for participants with a technical background. We strongly encourage reading the paper ahead of time to fully engage with the discussion.

This paper synthesizes findings from multiple recent studies to demonstrate that when Large Language Models show their step-by-step reasoning, these explanations frequently diverge from their actual computational processes. The authors document systematic patterns of unfaithfulness: models rationalize answers influenced by subtle prompt biases without mentioning these influences, silently correct errors in their reasoning chains while still reaching correct conclusions, and use memorized shortcuts while presenting elaborate logical derivations.

The paper goes beyond documenting these failures to explore why they occur, examining mechanistic evidence that transformers process information through distributed parallel pathways that clash with sequential verbal explanations. Drawing parallels to human cognitive biases like confabulation and post-hoc rationalization, the authors argue this gap may be inherent to current architectures. Join us to discuss their proposed directions for improvement - including causal validation methods that verify whether stated reasoning steps actually influence outputs, cognitive science-inspired error monitoring, and human oversight interfaces - and what this means for the future of interpretable AI systems we can genuinely trust.

The paper can be found here: https://www.alphaxiv.org/abs/2025.02v2

Location
Lorong AI (WeWork@22 Cross St.)
17 Went