

Daniel Beechey - Explaining Reinforcement Learning with Shapley Values: Theory and Algorithms
Reinforcement learning agents can achieve superhuman performance in complex decision-making tasks, yet they typically cannot explain their actions. If we cannot understand an agent's decisions, should we trust it with control? In this talk, I ask from first principles what it means to explain a reinforcement learning agent, identifying three aspects of interaction that require explanation: behaviour, outcomes, and predictions. I then show how we can use Shapley values to place these explanations on a principled, game-theoretic foundation (SVERL, ICML 2023), and how they can be approximated in practice (FastSVERL, NeurIPS 2025), making Shapley-based explanation a practical tool for understanding the agents we deploy.
Bio: Daniel Beechey is a researcher at H Company in London, using reinforcement learning to train computer-use agents. He was previously a research scientist on the AI Agents team at Huawei's Noah's Ark Lab, and completed his PhD on explainable reinforcement learning in the Bath Reinforcement Learning Lab.