

Beyond the Black Box: A Technical Introduction to Mechanistic Interpretability
We've built very capable AI used in high-stakes domains, but have spent far less effort learning to understand it well and read its mind.
This session is a working, technical introduction to mechanistic interpretability: the neuroscience of AI.
The session will cover why this matters and then go through the core toolkit of mechanistic interpretability, going deep on sparse autoencoders, probes, and others to find the model's tangled internal state, use it for prediction, and perform steering and other causal tests.
We put the tools to work.
From the LLM side: finding and moving a feature, turning a model's internal knowledge into a signal, and catching hallucinations from the inside. From the agentic side: reading a model's internal state to flag a risky tool call before it executes, not after.
The field is roughly where neuroscience was in its early days. The tools are open, and the interesting problems are unclaimed. This one is meant as much for inspiration as instruction.
About the speaker:
Hariom has years of experience bridging AI, machine learning, quantitative techniques, and finance.
He is an O'Reilly author and published researcher, with multiple research contributions in AI, machine learning, and mechanistic interpretability, particularly focused on making large language models more transparent and reliable in financial and agentic AI settings.
He has been a featured speaker at several conferences and industry forums and received the Indian Achiever Award in Machine Learning. He completed his MS at UC Berkeley and his BE at IIT (India).
To attend online:
Gmeet link: meet.google.com/fxk-wkoi-xja
Looking forward to seeing you!