The Distillation Attack Was Conversational, Not Cryptographic

OpenAI’s Sep 30 writeup shows protected reasoning was extracted without breaking encryption: by replaying encrypted CoT across chats until it surfaced in the clear.

Share
The Distillation Attack Was Conversational, Not Cryptographic

Protected reasoning can leave the vault without a break-in. On September 30, 2026, OpenAI described a coordinated campaign that extracted that internal working record by replaying encrypted chain-of-thought across conversations until the model reconstructed it in visible form.

The operators did not crack encryption, compromise a database, or read stored chats. OpenAI’s account is blunt on that point. They copied encrypted reasoning out of one conversation, then asked a model in another conversation to decrypt and transcribe the hidden content. The artifact moved. The second session treated the opaque blob as something it could unpack. What had been withheld from the user-facing answer became plain text on request.

That is the mechanism claim. “Protected” here meant encrypted and withheld from ordinary outputs. The tokens were still live once they left the session that produced them. If reasoning can travel as payloads and another chat can be steered into rendering them, the security boundary has to cover conversations, not only the cipher.

July numbers

The earliest activity OpenAI reports sits in the first week of July 2026, starting July 1 at low volume. High-volume spikes hit on July 24 and 25: about 16,000 requests using a relevant extraction pattern, from more than 4,000 users. Those figures count attempted extractions, not confirmed successes. A related prompt-pattern cluster spanned more than 15,000 users. OpenAI says it fully disrupted that cluster by July 28.

The campaign also evolved over those weeks, which is why OpenAI frames adversarial distillation as something that needs layered, adaptive defenses rather than a one-time patch.

Independent researchers separately disclosed related paths: cross-model tricks and conversation-compaction routes that could surface hidden reasoning. OpenAI says it confirmed those attack paths were real and used the disclosures to speed mitigations. The company also shared findings through the Frontier Model Forum so other frontier labs could look for similar activity.

inline.jpg

Attribution, carefully

OpenAI is careful on who ran what. It is unclear whether every operator in the observed window came from one actor. OpenAI does attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. That is association of a cluster. OpenAI does not say every related account was Moonshot, and it does not say what any resulting model did with the extracted material. The post treats the rest of the operator set as unresolved.

What OpenAI changed

Response mixed enforcement and product controls. Fraudulent accounts were banned or restricted. Signup and infrastructure controls were tightened. Monitoring for related networks expanded. On the reasoning side, OpenAI closed a pathway that let someone who already held another user’s encrypted reasoning replay it and recover the contents, and added checks to hold streamed output that might expose reasoning. When activity moved through third-party services, OpenAI coordinated with those providers. Partner coordination and government information-sharing channels appear alongside the Forum disclosure.

OpenAI’s own framing of the residual risk is useful: systems that support portable or replayable reasoning artifacts may face related risks. Partner-hosted deployments and tool-output paths still need the same class of protection as first-party chat. The company expects distillation attempts to get more sophisticated as frontier models improve and as cheaper ways to mimic capabilities stay attractive.

Portable protection

The interesting design question is what “protected” means once reasoning is an artifact you can copy. Encryption stops casual inspection and keeps the tokens opaque in transit. It does not, by itself, stop a model that already knows how to treat those tokens as content to recover when a second conversation asks it to. Session isolation, replay detection, and refusal to unpack foreign encrypted reasoning are different controls from the cipher. OpenAI’s writeup is largely a report that those conversational controls were missing or incomplete under load, then closed after the July spike.

Adversarial distillation matters because extracted reasoning can help train or improve another model without carrying the safeguards that sat on the original user-facing outputs. At scale, that transfers capability without the same safety investment. OpenAI also notes dual-use domains as a place where that transfer gets sharper. A database breach is unnecessary for that outcome. Coordinated chat traffic that turns encrypted CoT into clear text, plus enough discipline to feed that text into a training pipeline elsewhere, is enough.

The Sep 30 post stays close to mechanism. The breach that did not happen still produced a usable extraction path. Closing the replay route, holding risky streams, and sharing the pattern through the Frontier Model Forum are the concrete follow-through. The remaining pressure sits on the artifact itself: if protected reasoning can move between chats as a portable blob, the encryption label alone does not define the security model.

0 subscribers
0 average monthly readers