Cover Image for AI Safety Reading Group - Verbalizing Latent Behavior
Cover Image for AI Safety Reading Group - Verbalizing Latent Behavior
10 Went

AI Safety Reading Group - Verbalizing Latent Behavior

Hosted by Campbell Hutcheson & Era Qian
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Reading papers related to what is learned in fine tuning and LLMs ability to verbalize things on which they have been trained.

Tell me about yourself: LLMs are aware of their learned behaviors

https://arxiv.org/abs/2501.11120

Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data

https://arxiv.org/abs/2406.14546

Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

https://arxiv.org/abs/2407.04694

Location
717 Market St
San Francisco, CA 94103, USA
10 Went