AI

OpenAI Disrupts Distillation Campaign Linked to Moonshot AI

Published  ·  8 min read

OpenAI said on Wednesday that it identified and disrupted a coordinated distillation campaign designed to illicitly extract protected reasoning from its AI models, and the company attributed a core cluster of the activity to individuals associated with Moonshot AI, a Chinese AI company based in Beijing.

OpenAI did not cite any technical evidence to back that assessment, likely for security reasons, and that absence of public proof is worth noting because attribution in the AI space is often asserted without the kind of forensic detail you would see in a traditional malware report.

The company was clear about what did not happen, saying the operators did not break encryption, compromise a database, or gain direct access to stored user conversations, and instead they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated OpenAI's terms of service.

Quick Summary

What

Details

Campaign

Adversarial distillation

Attributed To

Individuals associated with Moonshot AI

Timeline

July 1 to July 28, 2026

Peak Activity

16,000 requests from 4,000 users on July 24-25

Related Activity

15,000+ users in prompt-pattern activity

OpenAI Response

Banned accounts, added mitigations, closed replay pathway

How the Campaign Unfolded

The activity is said to have begun on July 1, 2026, initially at a low volume before it spiked on July 24 and 25, 2026, to 16,000 attempted requests using a relevant extraction pattern from over 4,000 users.

Upon further investigation, OpenAI said it identified related prompt-pattern activity across more than 15,000 users, and the campaign was fully disrupted on July 28, 2026.

So the timeline is short, just under a month, but the scale is significant, because thousands of accounts were involved in what appears to have been a coordinated effort to pull reasoning traces out of the model at volume.

What Adversarial Distillation Actually Means

OpenAI characterized the activity as adversarial distillation, which involves the systematic and unauthorized use of one model's outputs to help train, reproduce, or improve another model.

That is a fancy way of saying someone is using your model to teach their model, and it is a problem because it lets a competitor skip the expensive work of building capabilities from scratch while also bypassing the safety measures you built into your own system.

OpenAI said it has since deployed additional mitigations to combat this attack and banned the fraudulent accounts engaged in the activity, and it also closed a pathway that made it possible for someone who already possessed another user's encrypted reasoning to replay it and recover its contents, alongside adding checks to detect and hold streamed output that might expose reasoning.

That last detail is important, because it suggests the attack was not just about volume, it was also about finding a technical loophole that let encrypted reasoning be replayed and decoded.

The Encrypted Reasoning Vulnerability

In a study published in August 2026, a group of researchers found an architectural vulnerability impacting Claude, Gemini, and GPT that made encrypted reasoning traces fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem.

An attacker could exploit this technicality to develop a scalable decryption jailbreak and circumvent anti-distillation mechanisms, and the researchers from MATS Research, ELLIS Institute Tübingen, and Synk described the technique in blunt terms, saying that by injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, they force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly.

That is a serious design flaw, because it means the encryption protecting reasoning traces is not actually a security boundary if the same provider hosts weaker models that can be tricked into decrypting them.

The study also noted that the technique allows for large-scale private data extraction, opens the door for invisible prompt injections by embedding malicious payloads entirely within encrypted blocks, and inadvertently reveals hazardous information hidden within the reasoning process, even if the model's final visible output rejects a harmful request.

So the risk is not just about stealing reasoning, it is also about smuggling malicious instructions and surfacing harmful content that the model would otherwise refuse to produce.

Why OpenAI Says This Matters

Given that protected reasoning offers insights into how a model works its way through a task, extracting this information can reveal sensitive data and help others reproduce the model's capabilities, according to OpenAI.

The company said adversarial distillation poses safety and national security risks, because extracted reasoning could be used to train another model without preserving the safeguards applied to the original model's user-facing outputs.

At scale, distillation can also accelerate the transfer of advanced capabilities without requiring the same investment in safety, and these concerns become heightened as models gain capabilities in dual-use domains.

That is a fair point, because the whole premise of AI safety work is that safeguards are built into the model itself, and if someone can extract the reasoning and train a new model without those safeguards, the safety work is effectively bypassed.

The Moonshot AI Accusations Are Not New

This is not the first time Moonshot AI has faced distillation accusations, because last month, rival Anthropic accused Moonshot AI of stealthily relaying customer requests to Claude as opposed to processing them using Kimi, and then displaying responses from Claude back to the users.

The company is also alleged to have retained a subset of these exchanges to train its chain-of-thought model, and the activity has been tracked under the moniker GTG-16002.

So there is a pattern here, or at least a pattern of accusations, and while OpenAI did not provide technical evidence for its attribution, the fact that two major AI labs have now made similar claims about the same company is notable.

What This Means for the AI Industry

The distillation problem is not going away, and it is not limited to one company or one country, because the economics of AI development create a strong incentive to extract capabilities from frontier models rather than build them from scratch.

That incentive is only going to grow as models become more capable and more expensive to train, and the technical loopholes that make distillation possible, like the encrypted reasoning replay issue, are likely to be discovered and exploited by more actors over time.

So the industry needs to treat reasoning traces as sensitive data, not just as internal artifacts, and it needs to design systems where the encryption actually holds across different models and different sessions, because right now that is not guaranteed.

What You Should Do

  • If you are an AI developer, review how your platform handles encrypted reasoning traces, because the replay vulnerability described in the August 2026 study affects multiple providers.
  • Monitor for coordinated extraction patterns, such as high volumes of requests from many accounts using similar prompts, because that is what this campaign looked like on the wire.
  • Treat reasoning traces as sensitive data, because they can reveal internal model behavior, training data, and safety mechanisms.
  • If you are an enterprise buyer, ask your AI vendors how they prevent distillation and what safeguards they have in place.
  • Stay informed about distillation research, because the techniques are evolving quickly and the vulnerabilities are not always obvious.
  • If you are a policy maker, consider that distillation has national security implications, because it can transfer advanced capabilities without transferring the safety work that goes with them.

The Bottom Line

OpenAI disrupted a coordinated distillation campaign that it links to individuals associated with Moonshot AI, and while the company did not provide technical evidence for the attribution, the campaign involved thousands of accounts and tens of thousands of requests, and it exploited a vulnerability in how encrypted reasoning traces are handled across models, so the incident is a reminder that AI capabilities can be extracted in ways that bypass the safeguards built into the original system.

Quick Reference

Key Point

Detail

Campaign

Adversarial distillation

Attributed To

Individuals associated with Moonshot AI

Timeline

July 1 to July 28, 2026

Peak

16,000 requests from 4,000 users

Related Activity

15,000+ users

Vulnerability

Encrypted reasoning replay across models

OpenAI Response

Banned accounts, closed pathway, added checks

What to Do

  • Review encrypted reasoning trace handling
  • Monitor for coordinated extraction patterns
  • Treat reasoning traces as sensitive data
  • Ask vendors about distillation safeguards
  • Stay informed about distillation research
  • Consider national security implications

FAQ Section

What is adversarial distillation?

It is the systematic and unauthorized use of one model's outputs to train, reproduce, or improve another model, which lets a competitor skip the expensive work of building capabilities from scratch.

Who did OpenAI attribute the campaign to?

OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI, a Chinese AI company based in Beijing, though it did not cite technical evidence for that assessment.

How large was the campaign?

The activity spiked to 16,000 attempted requests from over 4,000 users on July 24 and 25, 2026, and OpenAI identified related prompt-pattern activity across more than 15,000 users.

What vulnerability did the campaign exploit?

Researchers found that encrypted reasoning traces were fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem, allowing attackers to replay them and recover their contents.

What did OpenAI do about it?

OpenAI banned the fraudulent accounts, deployed additional mitigations, closed the pathway that allowed reasoning replay, and added checks to detect and hold streamed output that might expose reasoning.

Has Moonshot AI been accused of this before?

Yes, Anthropic accused Moonshot AI last month of relaying customer requests to Claude and then displaying responses from Claude back to users, and of retaining some exchanges to train its chain-of-thought model.

Source:
Professional Services

Explore Our Cybersecurity Services

Our insights are backed by hands-on service delivery. If your business needs professional cybersecurity support, our UK-based specialists are ready to help.

© 2016 – 2026 Red Secure Tech Ltd. Registered in England and Wales — Company No: 15581067