

A Discussion on Preferences, Personas, and What Drives LLM Behavior
In this interactive research discussion, Pranav Mahajan, a postdoctoral researcher at the University of Oxford, will present recent work investigating if and how preferences and personas impact behavior in LLMs. He’ll begin with his previous SPAR project on stated versus revealed preferences, which asks whether what models say they prefer corresponds to the choices they make under different elicitation procedures. He’ll then discuss his current work on how models’ beliefs about their users can shape behavior, including the model's persona selection.
The discussion will be hosted by Mohan Gupta, a postdoctoral researcher at Princeton University, whose group is exploring related questions about preference structures, task selection, and the mechanisms underlying LLM behavior.
We’ll use these projects as a starting point for a broader discussion about what kinds of explanations we should use for LLM behavior. In particular, we’ll consider questions such as:
What evidence would we need to say that an LLM has an underlying preference?
How should behavioral choices, self-reports, internal representations, and causal interventions relate to one another?
When does a “persona” provide a useful explanation, and when might lower-level mechanisms—such as representations, task selection, or control signals—be more informative?
The format will be approximately 20 minutes of presentation followed by an open research discussion.
Researchers and students interested in LLM behavior, AI alignment, interpretability, preferences, personas, emergent misalignment, or AI cognition are welcome.