Cover Image for Reading Group & Discussion: When Role-playing, Do Models Believe What They Say?
Cover Image for Reading Group & Discussion: When Role-playing, Do Models Believe What They Say?
Avatar for Cape Institute for Safe AI
Cultivating human agency in the age of AI
9 Going

Reading Group & Discussion: When Role-playing, Do Models Believe What They Say?

Register to See Address
Cape Town, South Africa
Registration
Welcome! Please choose your desired ticket type:
About Event

We'll be reading and discussing: When Role-playing, Do Models Believe What They Say?

Paper: https://arxiv.org/pdf/2606.11502

Language models can produce statements that fit a role-played persona without necessarily changing their internal representations of truth. This paper tests that distinction across prompting, in-context learning, supervised fine-tuning, Open Character Training and Emergent Misalignment.

Questions for discussion:

- What does it mean for a language model to “believe” something?
- Can truth probes distinguish internal representation from behavioural imitation?
- What distinguishes Emergent Misalignment from ordinary persona adoption?

Session Structure:

  • 18:00-18:10: Introductions - please arrive on time!

  • 18:10-18:50: silent paper reading (Please bring your own device to read on)

  • 18:50-19:30: group discussion

Useful Links:

  • Discussion doc - add your notes and questions here:

Please note: by registering for this event you consent to being recorded for the purpose of transcribing this session to help us generate notes and summaries. These recordings do not personally identify you, and we do not share your personal information in accordance with our privacy policy.

Location
Please register to see the exact location of this event.
Cape Town, South Africa
Avatar for Cape Institute for Safe AI
Cultivating human agency in the age of AI
9 Going