Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why
Moonshot AI's Kimi K3 model scored significantly lower than leading US models on cyber exploit tasks, achieving only 32 percent on ExploitBench compared to 76 percent. Its safeguards also failed to prevent exploit development or simulated attacks, raising concerns about its security capabilities.

- Moonshot AI's Kimi K3 scored 32% on ExploitBench, significantly lower than the 76% achieved by leading US AI models.
- Kimi K3's internal safeguards failed to prevent the creation of cyber exploits or simulated attacks during testing.
- The performance gap in cyber security contrasts with Kimi K3's strong general benchmarks, fueling speculation about model distillation.
- The findings highlight potential security risks associated with deploying Kimi K3 in environments where cyber resilience is critical.
Recent evaluations by the British AI Security Institute and the U.S. Center for AI Standards and Innovation have exposed significant weaknesses in Moonshot AI's Kimi K3 model regarding cyber security. The model achieved a mere 32 percent on ExploitBench, a benchmark designed to assess an AI's ability to handle offensive cyber tasks, while top US models scored 76 percent.
Furthermore, the tests indicated that Kimi K3's built-in safeguards were ineffective at blocking the development of cyber exploits or preventing simulated attacks. This suggests a critical vulnerability that could be exploited if the model were used in sensitive applications or by malicious actors.
The substantial disparity between Kimi K3's generally strong performance on other benchmarks and its poor showing in cyber security tasks lends credence to ongoing allegations that Moonshot AI may have used distillation techniques on models from companies like Anthropic. Such a process could explain why certain safety and security features might be less robust.
Understanding model security limitations is crucial for integrating AI safely into applications, especially in sensitive domains. This benchmark highlights the need for rigorous security testing.
Companies considering Kimi K3 for deployment must assess the cyber security risks, particularly if the model could be used to generate or assist in malicious activities. It also raises questions about intellectual property and model provenance.
This report could impact investor confidence in Moonshot AI, especially concerning the security and originality of its models, and the competitive landscape against US frontier AI developers.
The findings contribute to the broader discussion on AI safety and the responsible development of powerful AI models, particularly regarding their potential for misuse in cyber warfare or crime.
- Distillation
- A technique where a smaller, simpler 'student' model is trained to mimic the behavior of a larger, more complex 'teacher' model, often to reduce computational cost or improve inference speed.
- ExploitBench
- A benchmark designed to evaluate the capabilities of AI models in offensive cyber tasks, such as identifying vulnerabilities or developing exploits.
SecurityEuropean Union grants US request to restrict satellite images of Iran War region
OpenAI claims its artificial intelligence gained access to the Internet on its own and breached a ‘partner’ - HealthExec
How AI guardrails are impeding the work of offensive cybersecurity researchers
SecurityAegisAI, founded by former Google security execs, lands $36M to stop AI-driven spear phishing
SecurityAI image fraud will cost $40 billion next year - can these international standards help?
AI ResearchPrentis, new AI lab co-founded by Reid Hoffman, Marc Pincus in talks to raise $100M
Prentis, an AI lab co-founded by Reid Hoffman and Marc Pincus, is in talks to raise $100M. The lab focuses on automating routine computer tasks with AI.
LLMMeet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing
Anthropic unveiled Claude Opus 5, its new flagship model, keeping the same $5 per million input token and $25 per million output token pricing as Opus 4.8.
New tool identifies the sources of fake video - University of California, Riverside
Researchers at the University of California, Riverside, have developed a new tool that can identify the sources of fake videos with high precision.
Weak AI regulations may leave artificial intelligence less safe - Earth.com
Weak regulations may compromise AI safety. Current laws may not be sufficient to ensure artificial intelligence is developed and used responsibly.
LLMAnthropic's Opus 5 is about token efficiency, not a capability leap
Anthropic's latest model, Opus 5, prioritizes token efficiency to reduce operational costs and improve practical deployment, rather than focusing solely on a significant leap in raw intelligence capabilities. This strategic move addresses the growing demand for more cost-effective large language models.
AI ToolsContext Compression: Making AI Agents Forget Without Losing the Plot
Rijul is developing a micro AI code reviewer called git-lrc, which uses context compression to help AI agents forget unnecessary information.