Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

🎙 Machine Learning Street Talk 👥 219K 📅 August 22, 2026 ⏱ 49 min 👁 192 📄 expert opinion 🧭 2026-08-22
Available in: English (current) Français

Keywords

encrypted reasoningreplay attackjailbreakchain-of-thoughtmodel distillation

Summary

In this episode of Machine Learning Street Talk, Tim Scarfe interviews Ilia Shumailov and Alexander Panfilov about their paper ‘Stealing Reasoning Traces from Proprietary LLM APIs’. The researchers discovered a critical vulnerability in how frontier LLM providers (Anthropic, OpenAI, Google) handle encrypted reasoning traces. These traces, meant to be opaque to users, can be replayed across different user sessions and even across sibling models (e.g., from Opus to Haiku). By feeding a smaller model a reasoning trace from a larger model, the smaller model can be induced to decode and reveal the hidden reasoning in plain text. This enables a range of attacks: extracting private information from user sessions, performing prompt injections, and creating jailbreaks. The discussion covers the technical details of the attack, the portability of the encrypted blobs, the observation of ‘alien’ reasoning in some traces, and the potential for model distillation. The researchers also discuss responsible disclosure, the reactions from AI labs, and possible mitigations, including architectural changes and model-level defenses. They emphasize that while the vulnerability is serious, it is a specific instance of the broader jailbreaking problem, and they advocate for controlled experiments over sweeping claims.

191 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable first-hand insight into a significant security vulnerability. The authors clearly explain the attack mechanism, its implications, and the limitations of their findings. They argue convincingly that the vulnerability stems from a combination of architectural choices (replayable encrypted blobs) and model behavior (smaller models readily decoding traces). They also address potential counterarguments, such as the possibility of distillation evidence, and they are careful to distinguish between demonstrated attacks and speculative claims. The discussion is well-structured, moving from the basic bug to its broader safety implications and possible defenses.

Scientific Rigor, Source Quality, Title Accuracy

The discussion is grounded in the authors’ own research, which is available on arXiv. They reference several external sources, including a report from METR on model monitoring, an Anthropic research post on reasoning models, and a security incident involving OpenAI and Hugging Face. The title accurately reflects the content. The conversation is informal but the authors demonstrate a strong command of the subject matter. The description provides links to the paper and related resources, which adds to the overall credibility.

186 words

Title / Content Match

The title accurately reflects the core topic: the paper and its implications for proprietary LLM APIs.

Quality & Reliability

8/10

The discussion is led by the paper's authors, providing first-hand expertise. They describe a concrete vulnerability with reproducible methodology, and they reference external reports and incidents. However, the conversation is informal and lacks detailed technical depth, and some claims (e.g., model distillation evidence) are presented as preliminary observations.

Chapters

Cited Sources

Concurring Sources

Dissenting Sources

  • OpenAI and Hugging Face partner to address security incident during model evaluation — This incident is about a different security issue (data exposure during evaluation) and does not directly contradict the video's claims, but it shows that security incidents can have different causes.

External References

Contribution & Novelties

This video provides a detailed, first-hand account of a novel security vulnerability affecting all major LLM providers. The key novelty is the demonstration that encrypted reasoning traces can be replayed and decoded by smaller models, enabling a range of attacks. This goes beyond previous work on jailbreaking by showing a systemic architectural flaw. The discussion also highlights the potential for model distillation and the observation of ‘alien’ reasoning patterns in real-world traces.

Pour aller plus loin :

  • Chain-of-thought prompting — Background on the reasoning process that is being exposed.
  • Model extraction attack — Related concept of extracting model knowledge.
  • Jailbreaking (AI) — The broader category of attacks this vulnerability falls under.

111 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable discussion. The high 'quantite_information' and 'qualite_information' scores reflect the depth and relevance of the content, while the 'niveau_technique' score indicates a moderately technical level suitable for a general technical audience. The 'fiabilite_globale' score is high due to the authors' expertise and the inclusion of references.

Reliability 8/10