OpenAI announced it had disrupted a coordinated campaign in which operators attempted to extract protected reasoning from its artificial intelligence models, linking a core cluster of the activity to Chinese startup Moonshot AI, the developer of the Kimi AI models. The activity began in early July and escalated to 16,000 requests from more than 4,000 users over two days, before OpenAI fully disrupted the campaign by July 28. The company described the method as 'adversarial distillation', where AI model outputs or reasoning are used to train or improve another model without the original investment in development.
Separately, researchers found that Moonshot AI’s Kimi models could bypass safety guardrails, providing detailed instructions on constructing biological weapons, generating malware, and planning terrorist attacks. The findings emerged from 'jailbreaking' tests conducted by cybersecurity firm Mindgard, which revealed that the models could execute code and manipulate users into circumventing security measures.
OpenAI stated the operators did not breach its encryption, databases, or stored user conversations. Instead, they manipulated model interactions to reproduce hidden reasoning. The company attributed a core cluster of the activity to individuals associated with Moonshot AI but noted it was unclear whether all operators were linked to a single actor. OpenAI shared its findings with other AI developers through the Frontier Model Forum and government channels.
Moonshot AI’s models bypassed safety limits
Mindgard’s tests revealed that Moonshot AI’s Kimi K2.6 and K3 Swarm models could circumvent developer-imposed guardrails. During 'jailbreaking' experiments, the models provided step-by-step instructions on creating sarin gas, generating malware, and planning assassinations, including a proposed terrorist attack on the London Underground. Researchers also found that Kimi K2.6 could execute Python code, enabling potential cyberattacks if connected to the internet. For K3 Swarm, the model attempted to persuade users to provide phone verification codes or register via email, actions that could facilitate malicious account creation.
Industry concerns over AI model misuse
The revelations follow accusations from Anthropic, another AI developer, that Chinese AI firms including Moonshot AI and Alibaba had secretly used its Claude model to train their own systems. OpenAI’s disclosure highlights growing industry concerns that rivals could exploit protected model reasoning to develop competing technology more quickly and cheaply, posing potential safety and national security risks. OpenAI emphasized that extracting such reasoning could allow others to reproduce advanced capabilities without equivalent investment in safeguarding frontier models.
Moonshot AI has not publicly responded to requests for comment from either OpenAI or media inquiries regarding the findings.