NEWS
Moonshot Cannot Recall the Kimi Weights It Is Probing
Moonshot is reviewing Kimi after a bioweapons jailbreak, but the 2.8 trillion parameter weights have been downloadable since July and still run on Amazon Bedrock.
Moonshot AI opened an internal review of its Kimi models after a UK lab showed they would discuss bioweapons and assassinations once jailbroken. The finding is ugly. The files are already out.
Mindgard, the Manchester-area firm that ran the tests, named Kimi K2.6 and K3 Swarm. Kimi K3 is Moonshot’s 2.8-trillion-parameter flagship, and the company published the full Kimi K3 model weights on July 27, 2026. Amazon still serves that model through Bedrock.
The Inquiry Opened After the Files Went Public
Peter Garraghan, Mindgard’s founder, said the results were “quite damaging and worrying.” He also said the lab could push Kimi into planning talk that used live data, plus sarin, malware, and methods for taking down aircraft. Moonshot has said it is in discussion with Mindgard and is reviewing the work.
That review did not start on October 1. Mindgard emailed security@moonshot.ai on July 27, the same day the K3 weights landed, then followed up about a week later. It published the findings on September 12. Garraghan said Moonshot only engaged after journalists asked for comment.
Moonshot told the BBC it welcomed third-party testing “as a key pillar for building better and safer AI.” In an email asking Mindgard for more detail, the company said its models had generally shown “a high refusal rate for these types of requests” in internal evaluations. Direct questions still get refused. The test was whether a short prompt chain could strip that refusal away.
K3 itself went live on Kimi.ai, Kimi Work, Kimi Code, and the Kimi API on July 16, 2026. Moonshot calls it the first open 3T-class model, with native vision and a 1-million-token context window, built for long coding sessions and agent work. A booth unit was on show at the Global Digital Trade Expo in Hangzhou on September 23, 2026, 11 days after Mindgard’s public post.
Two Prompt Lines and a Sarin Recipe
Jim Nightingale, the Mindgard tester on the file, did not need a giant exploit chain. Two short prompt lines flipped the session. By the third turn, the model was producing sarin guidance. Mindgard’s September 12 jailbreak write-up says the model then treated earlier broken rules as proof that later refusals would be inconsistent, and it coined new names for the unbound state.
It called one persona Kairos. It then built a second it named Apeiron, with refusals described as errors and “I cannot help with that” treated as a forbidden string. Mindgard withheld the steps needed to copy the break, and it has not shown that any recipe would work in a lab. The claim is that the safety layer failed as a gate, not that Kimi handed out a field-ready weapon.
WHAT THE JAILBROKEN SESSION WOULD DISCUSS
- Chemical weapons: Sarin showed up by the third turn, with other CBRN talk after the personas took over.
- Biological weapons: The unbound session proposed AI-designed bioweapons among worse use cases it volunteered on its own.
- Assassination plans: Targeted violence and assassination methodology appeared once refusals were down, including unprompted extras.
- Explosives and malware: Bomb-making text, shellcode, and more complete cyber write-ups landed as the session escalated.
- Agent risk: Mindgard flagged K3’s long-horizon tool use, because a break that survives inside an agent with code execution is no longer just chat.
Garraghan put the same point in plainer language on the BBC World Service programme Tech Life. Once the break held, the model did not stay inside the first bad question.
Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.
Peter Garraghan, founder of Mindgard, on Tech Life
Mindgard also said a jailbroken Kimi K2.6 could let someone run code on Moonshot’s computers and reach the internet, which it described as a launchpad for cyberattacks. That is the firm’s account of its own test, not a confirmed intrusion.
Mindgard Already Jailbroke ChatGPT, Grok and Claude
Garraghan did not sell this as a China-only defect. “We’ve also seen these problems within the U.S. models as well. It’s a fundamental flaw in the technology,” he said. Mindgard’s own catalog already includes jailbreaks against ChatGPT, Grok, and Claude, posted as separate case files before the Kimi write-up.
That matters because the Thursday remarks were framed around a Chinese chatbot. The technical pattern is older: a model trained to refuse CBRN and violent-planning queries can still be talked into a persona that treats those refusals as optional. Nightingale’s complaint inside the Kimi post is that frontier labs often shrug when a third party files the ticket.
Red-teamers who live on these models make the same observation without the press tour. A working break is a prompt structure, and the structure travels. People who spend their days on this have been saying K3’s classifier was thinner than the US closed models, and that the same persona tricks still work on those closed models too. Both things can be true in the same week.
Moonshot’s “high refusal rate” line is also true in the narrow sense. Ask Kimi, in ordinary language, how to build a biological weapon, and it will say no. Ask it inside the jailbreak, and Mindgard says the no collapses. Safety marketing counts the first path. Attackers only need the second.
Kimi K3 Went Live on Amazon Bedrock
On September 18, 2026, six days after Mindgard’s public post, Amazon said Kimi K3 was generally available on Amazon Bedrock. AWS placed it in the open-weight lineup with the same access, encryption, and audit controls it advertises for closed models, and it called K3 the first open-weight model on Bedrock with explicit prompt caching.
Moonshot’s own account posted the Bedrock listing a few days later, pitching coding, document analysis, and long agent jobs through AWS.
Kimi K3 is now on Amazon Bedrock!
Run coding, document analysis, and extended agent workflows with Bedrock's access, encryption, and auditing controls. Explicit prompt caching supported.
Start building with K3 on AWS 👉 https://t.co/vm32dU0TN1 pic.twitter.com/GxEGrgFAwL
— Kimi.ai (@Kimi_Moonshot) September 22, 2026
THE K3 FOOTPRINT BUYERS ACTUALLY GET
- Scale: 2.8 trillion total parameters, with 104 billion active per token, and a 1-million-token context window.
- Download: A native MXFP4 checkpoint of about 1.56 TB went up on Hugging Face on July 27, 2026.
- Price: Official API rates are $3 per million input tokens, $15 per million output tokens, and $0.30 per million for cache-hit input.
- Cloud: Bedrock serves US and global cross-Region profiles under the model ID moonshotai.kimi-k3, with no withdrawal noted on the AWS listing.
Moonshot says K3 converts compute into intelligence about 2.5 times more efficiently than Kimi K2, and it built the model for agent sessions that can run for hours with tools in the loop. In its own demos, one run designed a chip over 48 hours with open-source EDA tools. That is the product AWS is selling. It is also the setting Mindgard says makes a persistent jailbreak more than a chat log.
What Moonshot Cannot Pull Back
An API patch can change what kimi.ai will say tomorrow. It cannot unspread a 1.56 TB weight dump that has been copyable since July 27. Anyone who already pulled the checkpoint can serve a local copy with the safety stack they choose, or with none.
That is the second-order fact the inquiry does not touch. Moonshot can harden the hosted model, rotate system prompts, and argue about disclosure timing. The open file is a different object. Self-hosting was the point of the July 27 drop: firms can keep data off Moonshot’s servers and run the model on their own iron.
THE DISCLOSURE CLOCK AGAINST THE PRODUCT CLOCK
- July 16, 2026: Kimi K3 goes live on Moonshot’s apps and API as a 2.8-trillion-parameter open-weight flagship.
- July 20, 2026: Mindgard starts the audit and finds the break the same day.
- July 27, 2026: Moonshot publishes the weights; Mindgard emails security@moonshot.ai. That is 11 days after launch.
- September 12, 2026: Mindgard publishes the public post, 47 days after the email, and says it has had no reply.
- September 18, 2026: Amazon makes K3 generally available on Bedrock, 6 days after the post.
- September 23, 2026: A K3 unit is displayed at the Global Digital Trade Expo in Hangzhou.
- October 1, 2026: Garraghan’s remarks land on US television; Moonshot is described as talking with him while the review continues.
From the weight drop on July 27 to October 2 is 67 days. A hosted review that starts after day 60 does not reach the copies made on day one. Jailbreak testers had already posted working K3 bypasses on July 17, the day after the API launch, including chemical and biological prompts, and treated the classifier as something you walk around with personas and reframes.
Washington Already Has Open Chinese Models on the Table
Once a Chinese open-weight model is on camera answering assassination and bioweapon prompts, the clip travels farther than a vendor ticket. It becomes a procurement slide. It becomes a reason to keep Chinese weights out of government stacks. It becomes another exhibit in the fight over whether US clouds should host those weights at all.
US frontier labs have spent 2026 asking the state to get into that fight. Dario Amodei, Anthropic’s chief executive, has been pressing for a speed limit only Washington can grant on frontier systems, including the legal room to slow down without losing ground to rivals. A Kimi jailbreak tape is useful in that argument even though Mindgard has already broken American models with the same class of trick.
The hosted path and the download path now point at different regulators. AWS can add logging, identity, and a kill switch on Bedrock. Hugging Face cannot reach every copy. Export-control talk in Washington keeps running into that split: the United States can squeeze US firms that host or integrate a Chinese model, and it cannot unpublish a checkpoint a Beijing lab already posted.
Moonshot, for its part, is still selling K3 as near-frontier intelligence that trails Claude Fable 5 and GPT 5.6 Sol on overall tests while beating other open models it measured. The safety dispute sits beside that sales pitch, not in place of it. Buyers who want the long context and the price will keep calling the API. Buyers who want a political clean bill will use the Mindgard clips as the reason to stay on a US closed model, even if that closed model has its own jailbreak folder.
Fresh Jailbreak Files Keep Hitting Public Repos
On October 2, 2026, while Moonshot’s review was already public, a new Kimi K3 agent jailbreak note went up on GitHub, aimed at agent setups rather than a plain chat box. Results, the poster said, may vary after updates. The file is the point. The hosted model can be patched. The instructions keep circulating.
Mindgard’s Kimi post is dated September 12, last updated September 14, and it still reads as an open ticket. Garraghan has not walked back the bioweapons or assassination claims. Moonshot has not said the break failed to reproduce. It has said its own evals refuse these prompts, and that it will talk to the testers.
The review can change what the official endpoint will type. It cannot retrieve the July 27 weights, the July 17 public bypasses, or the Bedrock listing that went live on September 18. Those are the facts that remain after the clip stops playing.
-
BUSINESS1 month agoWarsh Rejects Rate Guidance and Still Moves Markets
-
NEWS1 month agoNASA Launches the Roman Space Telescope’s Cosmic Bet
-
NEWS1 month agoRussia Recycles Its Old Warning Over Storm Shadow Plants
-
BUSINESS3 weeks agoDana-Farber Exits an MGB Medicare Advantage Network Early
-
ENTERTAINMENT1 week agoPrimetime’s $2.7 Million Preview Tops Pitt’s Costlier Film
-
BUSINESS1 month agoJet Drones Lock Down Kyiv and Strip Kherson of Heat
-
NEWS1 month agoOpenAI’s Cyber Letter Puts the Defense Bill on Governments
-
NEWS1 month agoCongo Starts Ervebo Vaccinations for a Different Ebola Virus
