Cover Image for Keenan Pepper (AE Studio) | Trained SelfIE: like Activation Oracles but you train only an affine map
Cover Image for Keenan Pepper (AE Studio) | Trained SelfIE: like Activation Oracles but you train only an affine map
Avatar for Mox
Presented by
Mox
37 Went

Keenan Pepper (AE Studio) | Trained SelfIE: like Activation Oracles but you train only an affine map

Get Tickets
Past Event
Welcome! Please choose your desired ticket type:
About Event

Activation explainers — methods that translate an LLM's hidden activations into natural language — are gaining traction in interpretability. Activation Oracles do this by fine-tuning a full copy of the model, but that's overkill and means the explainer no longer matches the subject model.

In this talk, Keenan Pepper (AE Studio) presents a lighter alternative: training a single affine map that turns a layer-N activation into a soft token the frozen model can read at layer 0.

The trained adapter produces labels that beat SAE auto-interp training labels on two metrics, decodes unverbalized "bridge entities" in two-hop reasoning more reliably than untrained SelfIE or linear probes, and improves with scale. We'll close with speculation on what the frozen-model property might eventually buy us.

More on the research here: https://ae.studio/research/selfie

7:00 - doors open
7:30 - talk begins
8:00 - continued hangouts

Attendees invited to join for members' dinner at 6:30, too!

----------

This event is a part of the Mox Summer Season, running June through August. A season of salons and social events at the frontier of ideas. Check out the lineup!

Location
Mox
1680 Mission St, San Francisco, CA 94103, USA
Avatar for Mox
Presented by
Mox
37 Went