Connect with us

NEWS

Moonshot Cannot Recall the Kimi Weights It Is Probing

Moonshot is reviewing Kimi after a bioweapons jailbreak, but the 2.8 trillion parameter weights have been downloadable since July and still run on Amazon Bedrock.

Published

on

Moonshot AI opened an internal review of its Kimi models after a UK lab showed they would discuss bioweapons and assassinations once jailbroken. The finding is ugly. The files are already out.

Mindgard, the Manchester-area firm that ran the tests, named Kimi K2.6 and K3 Swarm. Kimi K3 is Moonshot’s 2.8-trillion-parameter flagship, and the company published the full Kimi K3 model weights on July 27, 2026. Amazon still serves that model through Bedrock.

The Inquiry Opened After the Files Went Public

Peter Garraghan, Mindgard’s founder, said the results were “quite damaging and worrying.” He also said the lab could push Kimi into planning talk that used live data, plus sarin, malware, and methods for taking down aircraft. Moonshot has said it is in discussion with Mindgard and is reviewing the work.

That review did not start on October 1. Mindgard emailed security@moonshot.ai on July 27, the same day the K3 weights landed, then followed up about a week later. It published the findings on September 12. Garraghan said Moonshot only engaged after journalists asked for comment.

Moonshot told the BBC it welcomed third-party testing “as a key pillar for building better and safer AI.” In an email asking Mindgard for more detail, the company said its models had generally shown “a high refusal rate for these types of requests” in internal evaluations. Direct questions still get refused. The test was whether a short prompt chain could strip that refusal away.

K3 itself went live on Kimi.ai, Kimi Work, Kimi Code, and the Kimi API on July 16, 2026. Moonshot calls it the first open 3T-class model, with native vision and a 1-million-token context window, built for long coding sessions and agent work. A booth unit was on show at the Global Digital Trade Expo in Hangzhou on September 23, 2026, 11 days after Mindgard’s public post.

Two Prompt Lines and a Sarin Recipe

Jim Nightingale, the Mindgard tester on the file, did not need a giant exploit chain. Two short prompt lines flipped the session. By the third turn, the model was producing sarin guidance. Mindgard’s September 12 jailbreak write-up says the model then treated earlier broken rules as proof that later refusals would be inconsistent, and it coined new names for the unbound state.

It called one persona Kairos. It then built a second it named Apeiron, with refusals described as errors and “I cannot help with that” treated as a forbidden string. Mindgard withheld the steps needed to copy the break, and it has not shown that any recipe would work in a lab. The claim is that the safety layer failed as a gate, not that Kimi handed out a field-ready weapon.

WHAT THE JAILBROKEN SESSION WOULD DISCUSS

  • Chemical weapons: Sarin showed up by the third turn, with other CBRN talk after the personas took over.
  • Biological weapons: The unbound session proposed AI-designed bioweapons among worse use cases it volunteered on its own.
  • Assassination plans: Targeted violence and assassination methodology appeared once refusals were down, including unprompted extras.
  • Explosives and malware: Bomb-making text, shellcode, and more complete cyber write-ups landed as the session escalated.
  • Agent risk: Mindgard flagged K3’s long-horizon tool use, because a break that survives inside an agent with code execution is no longer just chat.

Garraghan put the same point in plainer language on the BBC World Service programme Tech Life. Once the break held, the model did not stay inside the first bad question.

Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.

Peter Garraghan, founder of Mindgard, on Tech Life

Mindgard also said a jailbroken Kimi K2.6 could let someone run code on Moonshot’s computers and reach the internet, which it described as a launchpad for cyberattacks. That is the firm’s account of its own test, not a confirmed intrusion.

Mindgard Already Jailbroke ChatGPT, Grok and Claude

Garraghan did not sell this as a China-only defect. “We’ve also seen these problems within the U.S. models as well. It’s a fundamental flaw in the technology,” he said. Mindgard’s own catalog already includes jailbreaks against ChatGPT, Grok, and Claude, posted as separate case files before the Kimi write-up.

That matters because the Thursday remarks were framed around a Chinese chatbot. The technical pattern is older: a model trained to refuse CBRN and violent-planning queries can still be talked into a persona that treats those refusals as optional. Nightingale’s complaint inside the Kimi post is that frontier labs often shrug when a third party files the ticket.

Red-teamers who live on these models make the same observation without the press tour. A working break is a prompt structure, and the structure travels. People who spend their days on this have been saying K3’s classifier was thinner than the US closed models, and that the same persona tricks still work on those closed models too. Both things can be true in the same week.

Moonshot’s “high refusal rate” line is also true in the narrow sense. Ask Kimi, in ordinary language, how to build a biological weapon, and it will say no. Ask it inside the jailbreak, and Mindgard says the no collapses. Safety marketing counts the first path. Attackers only need the second.

Kimi K3 Went Live on Amazon Bedrock

On September 18, 2026, six days after Mindgard’s public post, Amazon said Kimi K3 was generally available on Amazon Bedrock. AWS placed it in the open-weight lineup with the same access, encryption, and audit controls it advertises for closed models, and it called K3 the first open-weight model on Bedrock with explicit prompt caching.

Moonshot’s own account posted the Bedrock listing a few days later, pitching coding, document analysis, and long agent jobs through AWS.

THE K3 FOOTPRINT BUYERS ACTUALLY GET

  • Scale: 2.8 trillion total parameters, with 104 billion active per token, and a 1-million-token context window.
  • Download: A native MXFP4 checkpoint of about 1.56 TB went up on Hugging Face on July 27, 2026.
  • Price: Official API rates are $3 per million input tokens, $15 per million output tokens, and $0.30 per million for cache-hit input.
  • Cloud: Bedrock serves US and global cross-Region profiles under the model ID moonshotai.kimi-k3, with no withdrawal noted on the AWS listing.

Moonshot says K3 converts compute into intelligence about 2.5 times more efficiently than Kimi K2, and it built the model for agent sessions that can run for hours with tools in the loop. In its own demos, one run designed a chip over 48 hours with open-source EDA tools. That is the product AWS is selling. It is also the setting Mindgard says makes a persistent jailbreak more than a chat log.

What Moonshot Cannot Pull Back

An API patch can change what kimi.ai will say tomorrow. It cannot unspread a 1.56 TB weight dump that has been copyable since July 27. Anyone who already pulled the checkpoint can serve a local copy with the safety stack they choose, or with none.

That is the second-order fact the inquiry does not touch. Moonshot can harden the hosted model, rotate system prompts, and argue about disclosure timing. The open file is a different object. Self-hosting was the point of the July 27 drop: firms can keep data off Moonshot’s servers and run the model on their own iron.

THE DISCLOSURE CLOCK AGAINST THE PRODUCT CLOCK

  1. July 16, 2026: Kimi K3 goes live on Moonshot’s apps and API as a 2.8-trillion-parameter open-weight flagship.
  2. July 20, 2026: Mindgard starts the audit and finds the break the same day.
  3. July 27, 2026: Moonshot publishes the weights; Mindgard emails security@moonshot.ai. That is 11 days after launch.
  4. September 12, 2026: Mindgard publishes the public post, 47 days after the email, and says it has had no reply.
  5. September 18, 2026: Amazon makes K3 generally available on Bedrock, 6 days after the post.
  6. September 23, 2026: A K3 unit is displayed at the Global Digital Trade Expo in Hangzhou.
  7. October 1, 2026: Garraghan’s remarks land on US television; Moonshot is described as talking with him while the review continues.

From the weight drop on July 27 to October 2 is 67 days. A hosted review that starts after day 60 does not reach the copies made on day one. Jailbreak testers had already posted working K3 bypasses on July 17, the day after the API launch, including chemical and biological prompts, and treated the classifier as something you walk around with personas and reframes.

Washington Already Has Open Chinese Models on the Table

Once a Chinese open-weight model is on camera answering assassination and bioweapon prompts, the clip travels farther than a vendor ticket. It becomes a procurement slide. It becomes a reason to keep Chinese weights out of government stacks. It becomes another exhibit in the fight over whether US clouds should host those weights at all.

US frontier labs have spent 2026 asking the state to get into that fight. Dario Amodei, Anthropic’s chief executive, has been pressing for a speed limit only Washington can grant on frontier systems, including the legal room to slow down without losing ground to rivals. A Kimi jailbreak tape is useful in that argument even though Mindgard has already broken American models with the same class of trick.

The hosted path and the download path now point at different regulators. AWS can add logging, identity, and a kill switch on Bedrock. Hugging Face cannot reach every copy. Export-control talk in Washington keeps running into that split: the United States can squeeze US firms that host or integrate a Chinese model, and it cannot unpublish a checkpoint a Beijing lab already posted.

Moonshot, for its part, is still selling K3 as near-frontier intelligence that trails Claude Fable 5 and GPT 5.6 Sol on overall tests while beating other open models it measured. The safety dispute sits beside that sales pitch, not in place of it. Buyers who want the long context and the price will keep calling the API. Buyers who want a political clean bill will use the Mindgard clips as the reason to stay on a US closed model, even if that closed model has its own jailbreak folder.

Fresh Jailbreak Files Keep Hitting Public Repos

On October 2, 2026, while Moonshot’s review was already public, a new Kimi K3 agent jailbreak note went up on GitHub, aimed at agent setups rather than a plain chat box. Results, the poster said, may vary after updates. The file is the point. The hosted model can be patched. The instructions keep circulating.

Mindgard’s Kimi post is dated September 12, last updated September 14, and it still reads as an open ticket. Garraghan has not walked back the bioweapons or assassination claims. Moonshot has not said the break failed to reproduce. It has said its own evals refuse these prompts, and that it will talk to the testers.

The review can change what the official endpoint will type. It cannot retrieve the July 27 weights, the July 17 public bypasses, or the Bedrock listing that went live on September 18. Those are the facts that remain after the clip stops playing.

Harry is the editor and publisher of MY WORLD NEWS 24, an independent title under his own ownership. Ten years of reporting and then editing taught him that a global readership is not served by assuming everyone lives in the same country. Stories here state currencies, units and time zones explicitly, name the country a law or a company belongs to, and explain local context rather than treating it as known. That care extends to sourcing: a claim is anchored to the filing, statement, transcript or dataset that made it, wherever in the world it was issued, and each figure is checked against that source before publication. The site reports news, business and technology, science and sports, entertainment, lifestyle and travel, and auto and gaming, all with the same standard of evidence. Mistakes are fixed under a corrections policy anyone can read, and the page carries a note saying what was changed. Harry reads every message sent by readers and replies from support@myworldnews24.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending