

Keenan Pepper (AE Studio) | Trained SelfIE: like Activation Oracles but you train only an affine map
Activation explainers — methods that translate an LLM's hidden activations into natural language — are gaining traction in interpretability. Activation Oracles do this by fine-tuning a full copy of the model, but that's overkill and means the explainer no longer matches the subject model.
In this talk, Keenan Pepper (AE Studio) presents a lighter alternative: training a single affine map that turns a layer-N activation into a soft token the frozen model can read at layer 0.
The trained adapter produces labels that beat SAE auto-interp training labels on two metrics, decodes unverbalized "bridge entities" in two-hop reasoning more reliably than untrained SelfIE or linear probes, and improves with scale. We'll close with speculation on what the frozen-model property might eventually buy us.
More on the research here: https://ae.studio/research/selfie
7:00 - doors open
7:30 - talk begins
8:00 - continued hangouts
Attendees invited to join for members' dinner at 6:30, too!
----------
This event is a part of the Mox Summer Season, running June through August. A season of salons and social events at the frontier of ideas. Check out the lineup!