Cover Image for How do we solve alignment? - Jason Brown
Cover Image for How do we solve alignment? - Jason Brown
Avatar for Meridian
Presented by
Meridian
21 Went

How do we solve alignment? - Jason Brown

Register to See Address
Cambridge, United Kingdom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

How Do We Solve Alignment?

This week we have Jason Brown, an Astra fellow and a regular of the Meridian community. From the man himself:

The talk will cover Edward and I's recent paper, "Understanding Goal Generalisation in Sequential Reinforcement Learning", the broader research agenda of Developmental Cognitive Interpretability (DCI), and how we might apply it to better understand generalisation in LLMs.

The paper focuses on analysing out-of-distribution goal-driven behaviour in maze-solving CNN agents. We examine how this OOD behaviour is dependent on which goal(s) were used to train the agent, and find several interesting phenomena. These include certain goal-features influencing goal generalisation more strongly than others, effects on OOD behaviour from early-in-training goals persisting, and early-in-training goals affecting how later-in-training goals are generalised from. We then introduce a mathematical framework and specific method for predicting an agents OOD behaviour purely from the specification of it's training pipeline, by observing how similar agents have generalised.

Stepping back from the paper, in our LW post on DCI, we lay-out the high-level ideas and assumptions that went into our work in the maze environment. Then we discuss the extent to which we think these assumptions hold in the LLM case, as well as lay out some concrete specific research directions and next-steps.

Location
Please register to see the exact location of this event.
Cambridge, United Kingdom
Avatar for Meridian
Presented by
Meridian
21 Went