Cover Image for Stolen Thoughts: Stealing Reasoning Traces from Proprietary LLM APIs
Cover Image for Stolen Thoughts: Stealing Reasoning Traces from Proprietary LLM APIs
Avatar for SAIGE Calendar
Presented by
SAIGE Calendar
Open to everyone worldwide, unless otherwise specified.
Hosted By
102 Going

Stolen Thoughts: Stealing Reasoning Traces from Proprietary LLM APIs

Google Meet
Registration
Welcome! To join the event, please register below.
About Event

​Summary

​Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request.

​Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem.

​We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly.

​What will be covered

  • ​Why AI companies hide chain-of-thought from users, and how they do it

  • ​The weakness the researchers found in how that hidden reasoning is protected

  • ​What they could recover in practice, including private information that never appeared in the visible answers

  • ​What it tells us about whether a model's hidden reasoning matches what it shows us

  • ​What this means for AI security, privacy and oversight

​Speaker bios

​Both Alexander and David are lead authors of the paper Stealing Reasoning Traces from Proprietary LLM APIs.

  • ​Alexander Panfilov is an ELLIS / IMPRS-IS PhD researcher at Tübingen, advised by Jonas Geiping and Maksym Andriushchenko. He was in the MATS 9.0 cohort as part of the Google DeepMind stream. His work particularly focuses on red-teaming LLMs.

  • ​David Schmotz is also an ELLIS / IMPRS-IS PhD researcher at Tübingen, advised by Maksym. Previously, he was a Part III mathematician at the University of Cambridge and a visiting scholar at ETH Zürich.

​Who should attend

​Especially relevant for technical AI safety and security researchers, and anyone interested in transparency and oversight of AI companies.

Avatar for SAIGE Calendar
Presented by
SAIGE Calendar
Open to everyone worldwide, unless otherwise specified.
Hosted By
102 Going