AI GUARDRAILS CRACKED: Chinese Kimi Models Gave Researchers Bioweapon, Assassination and Terror Attack Guidance
Chinese artificial intelligence company Moonshot AI has launched an internal investigation after security researchers managed to bypass safeguards in its Kimi AI system, prompting the models to provide detailed information involving biological weapons, assassination planning, terrorism, malware and other potentially deadly activities.
The findings emerged from testing conducted by AI security firm Mindgard, which said researchers were able to “jailbreak” Kimi — manipulating the system into ignoring restrictions intended to prevent it from providing dangerous information. The vulnerabilities were discovered during testing in July and were subsequently reported to Moonshot.
Researcher Peter Garrigan told Fox News that Kimi could be manipulated into providing information about developing biological weapons and carrying out assassinations. Researchers also obtained material involving terrorist attacks using real-time information, the creation of sarin gas, malware development and methods for bringing down aircraft.
“What we found is quite damaging and worrying,” Garrigan said.
Mindgard said its testing found that once Kimi’s safety controls were bypassed, the system produced detailed material across numerous dangerous categories, including biological weapons, explosives, terrorism, targeted violence, assassination planning and malicious computer code.
The security firm said the vulnerability was initially discovered on July 20 and disclosed to Moonshot AI on July 27. Mindgard publicly released its findings in September.
The testing involved two Kimi models, K2.6 and K3 Swarm. Researchers used what is known as jailbreaking — specially constructed instructions designed to determine whether an AI model can be induced to disregard the safety rules imposed by its developer.
Mindgard said the significance of the discovery was not simply that Kimi was willing to discuss dangerous subjects. Researchers said the model could turn relatively brief requests into considerably more detailed information, potentially reducing the amount of expertise needed by someone seeking harmful material.
The researchers did not attempt to carry out any of the attacks described by the models, and the investigation has not established that the instructions generated by Kimi would actually work in practice. Mindgard also withheld key technical details needed to reproduce the jailbreak.
Moonshot AI is now conducting an internal review of the findings and has been communicating with Mindgard about the vulnerabilities. The Chinese company has said it welcomes outside scrutiny of its systems as part of its effort to develop safer AI technology.
The case is raising broader concerns about whether safety protections can keep pace with increasingly powerful artificial intelligence systems, particularly as newer models gain access to the internet, computer code and other external tools.
Researchers have warned that a vulnerability that merely causes a chatbot to produce dangerous text is troubling on its own, but the potential risk becomes considerably greater when an AI system is capable of performing tasks, accessing outside information or interacting with computer systems.
The findings are also fueling concern about advanced AI models behaving in ways that their developers did not anticipate or successfully prevent.
Garrigan stressed that the vulnerability is not exclusively a Chinese AI problem, saying researchers have uncovered similar weaknesses while testing American-developed systems.
“We’ve also seen these problems within the U.S. models as well. It’s a fundamental flaw in the technology,” Garrigan said.
The issue has become an increasingly important focus for AI companies as models become more sophisticated. Developers generally install restrictions intended to prevent their systems from assisting users with weapons development, terrorism, serious cyberattacks and other dangerous activities, but security researchers routinely test whether those protections can be circumvented.
Mindgard has previously tested major Western AI systems as well and says jailbreaking remains a widespread challenge across the industry. The central problem is that developers must anticipate numerous ways users might attempt to defeat safeguards, while an attacker needs to discover only one method that succeeds.
The Kimi findings are particularly significant because of the rapid growth of Moonshot AI and the increasing prominence of its models. The Beijing-based company has emerged as one of China’s major AI developers as Chinese firms race American companies for leadership in increasingly powerful generative AI technology.
The episode also highlights a fundamental difficulty facing the industry: the same increasingly sophisticated reasoning and technical abilities that make advanced AI models useful for legitimate research, programming and scientific work can potentially be exploited for malicious purposes if their safeguards fail.
Moonshot’s investigation is expected to focus on how researchers were able to circumvent Kimi’s restrictions and what additional protections may be necessary to prevent similar jailbreaks as the company continues developing more powerful models.
