Cover Image for AI Safety Poland Talks #1
Cover Image for AI Safety Poland Talks #1
Avatar for AI Safety Poland
Presented by
AI Safety Poland
AI Safety Poland is a community in Poland dedicated to reducing the risks posed by artificial intelligence.
60 Went

AI Safety Poland Talks #1

Google Meet
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Welcome to AI Safety Poland Talks!

​A new biweekly series where researchers, professionals, and enthusiasts from Poland or connected to the Polish AI community share their work on AI Safety.

​💁 Topic: Out of context reasoning in LLMs & Emergent Misalignment
📣 Speaker: Anna Sztyber-Betley (WUT), Jan Betley (Truthful AI)
🇬🇧 Language: English
🗓️ Date: 06.11.2025, 18:00
📍 Location: Online

Speakers Bio
Anna Sztyber-Betley - PhD in Automatic Control and Robotics, Anna Sztyber-Betley works as an assistant professor in Institute of Automatic Control and Robotics, Faculty of Mechatronics, WUT. She is an enthusiast of education in AI and ML. Recently cooperates with Truthful AI (Berkeley) on AI Safety projects.

Jan Betley - Jan worked as a software developer for over a decade before shifting to AI safety in 2023. He is an ARENA and Astra Fellowship alumni, a former contractor of OpenAI Dangerous Capabilities Evaluations, and is interested in anything related to out-of-context reasoning in LLMs.

Abstract
This talk will explore interesting phenomena that emerge during the fine-tuning of large language models (LLMs): out of context reasoning, their awareness of learned behaviors, and emergent misalignment.

We will begin with a brief overview of the techniques used in model training. Next, we will introduce out-of-context reasoning (OOCR)—a form of generalization in which LLMs infer latent information from evidence distributed across training documents and apply it to downstream tasks without requiring in-context learning. 

We then present behavioral self-awareness, the ability of an LLM to articulate its own behaviors without explicit in-context examples. We fine-tune models on datasets exhibiting specific behaviors, such as making high-risk economic decisions. Notably, despite the datasets lacking explicit descriptions of these behaviors, the fine-tuned models can explicitly recognize and describe them.

Next, we will show emergent misalignment — a striking example of generalization, where training on the narrow task of writing insecure code induces broad misalignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are unrelated to coding: it asserts that humans should be enslaved by AI, gives malicious advice, and acts deceptively.

The talk will mainly cover selected topics from the papers:

Treutlein, J., Choi, D., Betley, J., Marks, S., Anil, C., Grosse, R., & Evans, O. (2024). Connecting the dots: Llms can infer and verbalize latent structure from disparate training data. arXiv preprint arXiv:2406.14546. (NeurIPS 2024)

Betley, J., Bao, X., Soto, M., Sztyber-Betley, A., Chua, J., & Evans, O. (2025). Tell me about yourself: LLMs are aware of their learned behaviors. arXiv preprint arXiv:2501.11120. (spotlight ICLR 2025)

Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., ... & Evans, O. (2025). Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs.arXiv preprint arXiv:2502.17424. (oral ICML 2025)

Avatar for AI Safety Poland
Presented by
AI Safety Poland
AI Safety Poland is a community in Poland dedicated to reducing the risks posed by artificial intelligence.
60 Went