

Alexander Panfilov - Stolen Thoughts: Stealing Reasoning Traces from Proprietary LLM APIs
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly.
Bio:
Alexander Panfilov is a third-year PhD student at ELLIS and the Max Planck Institute for Intelligent Systems in Tübingen, advised by Jonas Geiping and Maksym Andriushchenko. His research focuses on AI safety, particularly red-teaming LLMs, jailbreaks, and prompt injections. He was also a MATS 9.0 scholar in the Google DeepMind stream, where he worked on AI control.