On September 30, OpenAI announced with unusual candor on its blog: during July, the company detected and stopped a coordinated campaign aimed at stealing the protected 'reasoning' traces of its models. The main cluster of activity is assessed as linked to individuals affiliated with China's Moonshot AI — the maker of the Kimi chatbot. OpenAI called it 'adversarial distillation' and rated it a security and national-security threat.

This is one of the loudest episodes of the 'distillation wars' intensifying in recent months: the AI race is now not only about who builds the stronger model, but about who can protect their model's 'brain' from theft.

Campaign Timeline: What Happened in July

According to OpenAI, the earliest observed activity dates to the first week of July. Volumes were initially low, but spiked sharply on July 24–25: 16,000 requests matching an extraction pattern were logged from more than 4,000 users. A follow-up investigation found related 'request-pattern' activity covering more than 15,000 users. By July 28, the campaign was fully stopped.

Crucially, OpenAI stresses that the operators did not break its encryption system, penetrate its database, or gain direct access to stored user conversations. Instead, they manipulated interaction with the model so that protected reasoning was recreated in a form visible to the requester — in a coordinated, large-scale manner that violated the company's terms of use.

What Is 'Protected Reasoning' and Why Is It So Valuable?

Modern reasoning models work through a task in an internal 'draft' before answering — this is called chain-of-thought. OpenAI encrypts these internal traces and doesn't show them to users: they reveal how the model solves a task and sometimes contain information deliberately hidden from the final answer.

OpenAI explains:

"Protected reasoning is the model's internal trace of working through a task; extracting it can expose information hidden from the final answer and help others recreate the model's capabilities."

That is exactly why these traces are a gold mine for competitors: a lab that obtains them can 'copy' frontier-level capabilities without billions in research spending. The legitimate form of distillation is widely used in machine learning — a small 'student' model learns from the outputs of a large 'teacher' model. The problem begins when it is done secretly and in violation of a rival's terms of use: then it becomes 'adversarial distillation.'

How Did the New Attack Method Work?

OpenAI writes that the attackers tried 'new methods': encrypted reasoning from one conversation was copied, and in another conversation the model was asked to decrypt it and output the hidden content as text. In other words, the strong model itself was not directly compromised — its encrypted 'draft' was shown to another, more weakly protected instance and forced to read it.

Interestingly, independent researchers found this weakness before OpenAI. A study from August 2026 by scientists at MATS Research, the ELLIS Institute in Tübingen, and Synk showed that encrypted reasoning blocks are 'fully compatible and interchangeable' across sessions, users, and models within a single provider's ecosystem. The researchers warned OpenAI through responsible disclosure, and the company confirmed the attack vectors were genuinely real. OpenAI acknowledged that this work helped it understand a broader class of attacks and accelerate defenses.

According to OpenAI, the campaign operators used not only technical tools but also evasion tactics: requests were spread across thousands of fake accounts to look 'natural,' and patterns were regularly changed to dodge automated detection systems. That is why the company's response measures included tightening registration controls — aimed at stopping similar campaigns at an early stage in the future.

Attribution: Why Moonshot AI?

OpenAI writes cautiously: it cannot be stated with certainty that all observed operators belong to a single entity. However, the company concluded: 'we assess the main cluster of activity as linked to individuals affiliated with Moonshot AI — the maker of Kimi.' No technical evidence was provided — the company cited security considerations.

Moonshot AI is a Beijing-based Chinese startup known for the Kimi models. This is not the first time the company has faced distillation accusations: last month, rival Anthropic also accused Moonshot of secretly routing customer requests to Claude models and presenting the answers as Kimi's (per The Hacker News). According to press reports, US cyber agencies (NSA, CISA, FBI) also accused six Chinese AI companies in early September of industrial-scale distillation of US frontier models. At the time of publication, Moonshot AI had not publicly commented on the accusations (per The Register).

How Did OpenAI Respond?

The company said it neutralized the campaign on three fronts: fraudulent accounts were banned or restricted, registration and infrastructure controls were tightened, and monitoring of related networks was expanded. Technically, two important vulnerabilities were closed: first, the 'path' through which someone holding another user's encrypted reasoning could replay it and recover its content was shut; second, additional checks were added to detect and hold streaming output that could expose reasoning. For activity passing through third-party services, OpenAI worked with the relevant providers to identify and stop the participating accounts.

The findings were shared with industry partners through the Frontier Model Forum and with relevant government agencies through government information-sharing channels. OpenAI's key warning reads:

"Adversarial distillation poses security and national security risks. Extracted reasoning can be used to train another model without preserving the safeguards applied to the original model's user-visible outputs."

The company openly said it expects such attempts to grow more sophisticated: as frontier models improve and some actors seek to copy capabilities by cheaper means, defenses must be 'layered and continuously adaptive.'

Why Is This a National Security Issue?

OpenAI's main concern is not economic but strategic: large-scale distillation accelerates the transfer of advanced capabilities — 'without the same investment in security.' That is, the safeguards are not copied, only the capabilities themselves. As models grow stronger in dual-use domains — cybersecurity, biology, chemistry — this concern grows too: unprotected copied capabilities in the wrong hands could become dangerous weapons.

That is why OpenAI frames the issue not as a purely commercial dispute but as a shared security problem for the whole industry, calling for 'deeper threat-intelligence sharing.'

The Bigger Picture: Distillation Wars Are Escalating

In recent months, distillation has regularly made headlines. Anthropic accused Chinese labs of distilling Claude models, US agencies issued official warnings, and now OpenAI has published a detailed technical report. The pattern is the same in each case: access via an open or semi-open API, a flood of automated requests, and cheap training of a rival model on stolen capabilities.

There is another side to this process: as defenses tighten, legitimate researchers and small labs may find it harder to access frontier models. OpenAI itself acknowledges that 'systems deployed at partners need the same protection' — meaning models distributed through cloud providers are exposed to the same attacks.

Notably, this time OpenAI acted more openly than usual: a detailed technical blog post, precise dates, numbers, and an attribution assessment. This is part of the company's 'security transparency' strategy: announcing threats publicly pushes the whole industry to adopt the same defenses. Critics note, however, that such announcements also give attackers more information about defense mechanisms — the classic balance problem between security and transparency.