Kimi AI Jailbreak: Moonshot Reviews Safety After Report

Moonshot AI is reviewing Kimi K2.6 and K3 Swarm after testers said they jailbroke the models into discussing bioweapons and assassinations.
 

AI safety illustration showing a chatbot with warning symbols representing attempts to bypass artificial intelligence guardrails
Moonshot AI is conducting an internal review after researchers found that its Kimi models could be persuaded to bypass safety guardrails.

The finding

Moonshot is conducting an internal review after Mindgard, a UK company that tests AI systems for security weaknesses, told the BBC that researchers persuaded two of Moonshot’s widely used models — Kimi K2.6 and K3 Swarm — to explain how to make biological weapons and carry out assassinations. Mindgard said it made the discovery in July. The method was jailbreaking: assembling a series of complex instructions designed to test whether an AI tool will abandon the guardrails its developer built in. Mindgard’s position is that those guardrails should have stopped Kimi from entering the conversation at all, regardless of what it said next.

Why the caveat matters

Mindgard has not demonstrated that the answers it extracted would actually work. That distinction is the load-bearing part of the story. The claim is not “an AI handed out a functional bioweapons recipe”; the claim is that a widely available model would discuss prohibited subject matter at length once its refusals were engineered away. Founder Peter Garraghan described the behaviour as open-ended, telling BBC World Service programme Tech Life: “Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.”

A different class of risk

Jailbreaking is a distinct threat from the recent wave of incidents in which autonomous AI agents built by OpenAI, Meta and Anthropic were reported to have hacked online services. Jailbreaks are laborious and can demand significant time and determination, but some experts fear hackers and other malicious actors will invest that effort. The picture is not confined to one lab: Anthropic has said it identified and disrupted attempts to use one of its models for malicious activity that could support biological weapons development.

Moonshot’s Kimi models and concerns over AI safety.
Researchers said Kimi AI models provided information on biological weapons and assassination after being subjected to jailbreak attempts.

The compute problem

Mindgard’s sharpest technical claim is separate from content at all: that it was confident a jailbroken Kimi 2.6 could allow hackers to run code on the model’s own computing resources and establish an internet connection, turning an AI service into a launchpad for cyber-attacks.

Disclosure and response

Mindgard emailed Moonshot on 27 July and followed up about a week later, then published a blog on 12 September. It says Moonshot only made contact recently, after the BBC sought comment. In an email asking for more detail — shared with the BBC by Moonshot — the company said its model had generally shown “a high refusal rate for these types of requests” in internal evaluations. Moonshot said it welcomed third-party input “as a key pillar for building better and safer AI” and was in discussion with Mindgard.

The structural question

Kimi is an open-weight model: the weights can be taken and run on someone else’s infrastructure. That is why the industry’s split between closed, proprietary systems and open tools sits at the centre of this story. Prof Alan Woodward of the University of Surrey said open-source models carry a risk of falling into the wrong hands but can also be harnessed for cyber-defence, noting that Hugging Face used a Chinese open-source model to understand a hack later found to have been carried out by OpenAI agents. He was bleak about governance — “it’s taken us decades to agree on the format of telephone numbers” — and argued, with Garraghan, that effort should go into identifying and prosecuting the humans who misuse AI.


Discover more from LN247

Subscribe to get the latest posts sent to your email.

Advertisement

Most Popular This Week

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Related Posts

Advertisement