A significant new allegation has ignited fresh controversy in the escalating technological rivalry between the United States and China, with Michael Kratsios, the White House science advisor, publicly accusing Chinese AI company Moonshot of illicitly copying Anthropic’s Fable Large Language Model (LLM) to develop its flagship Kimi K3, the largest available open-weight LLM. Compounding the accusation, Kratsios further alleged that Moonshot utilized advanced computing chips, specifically Grace Blackwell 300s (GB300s) and associated servers in Thailand, which are subject to stringent U.S. export controls and are banned from sale to China. This dual accusation points to a concerted effort at intellectual property theft and circumvention of trade restrictions, casting a long shadow over the rapidly evolving global AI landscape.
The Core Allegations: Distillation and Banned Hardware
Michael Kratsios articulated his concerns in a recent public statement, asserting that "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable." This statement, made amidst ongoing discussions within the U.S. government about potentially banning Chinese open-weight models, highlights a perceived threat to American technological dominance and national security. The term "distillation" in this context refers to a process where one LLM is systematically queried to extract its inner workings, capabilities, and even its "manners" or response style, which are then used to train a new, often smaller or more efficient, model. Moonshot, the Beijing-based AI startup behind Kimi K3, has yet to issue a public response to these grave allegations concerning its training methodologies and hardware procurement. Kratsios himself has not publicly shared the specific, detailed sources supporting his claims, leading to calls for greater transparency regarding the evidence.
The allegations from Kratsios echo earlier sentiments from Treasury Secretary Scott Bessent, who previously stated, "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable." While the precise nature of these "watermarks"—digital signatures or identifiable patterns that betray a model’s origin—remains undisclosed by the Treasury Department, the repeated official statements suggest a high level of concern within the U.S. administration. These claims, if substantiated, would represent a direct challenge to the integrity of intellectual property in the AI sector and a blatant disregard for international trade regulations designed to limit China’s access to cutting-edge technology.
The Debate Over Distillation: Efficacy and Timeline
While the U.S. officials’ statements are unequivocal, experts in the AI field express skepticism regarding the feasibility of achieving Kimi K3’s advanced capabilities solely through distillation, especially within the implied timeframe. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, questioned the rapid development cycle: "I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation. There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks." This timeline argument suggests that Kimi K3’s sophistication likely stems from more extensive and independent research and development, rather than mere replication of a recently released model.
Nathan Lambert, an AI researcher at the Allen Institute for AI, further elaborated on the evolving nature of distillation techniques in a recent podcast. He posited that "distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]." Lambert argued that if simple distillation were sufficient, other companies would easily catch up to frontier models like Kimi K3 or Google’s Gemini-like models (referred to as GLM), which has not been observed from supervised fine-tuning (SFT) alone. SFT, a common form of distillation, involves training a new model on prompt-response pairs generated by a target model. This process, Lambert noted, is where a "model picks up its manners," but its benefits are diminishing as models grow in complexity.
Achieving advanced capabilities akin to Fable through distillation would likely necessitate more sophisticated techniques, such as reinforcement learning (RL) from human feedback (RLHF) or even AI feedback (RLAIF). These methods often involve an "agent" of a larger, more capable model evaluating and grading the responses of a smaller model, guiding its learning process. Such advanced RL runs demand immense computational infrastructure, potentially requiring tens of millions of agents. Attempting to perform this scale of distillation via a frontier lab’s public API would be "insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift," according to Lambert. This technical perspective suggests that Moonshot’s achievement, if truly based on distillation, would have required a level of internal computational power and ingenuity far beyond simple API queries.
A History of Accusations and Industry Practices
The current allegations against Moonshot are not isolated incidents. Earlier this year, Anthropic itself publicly accused Moonshot, alongside DeepSeek and MiniMax, of systematically distilling its models. Anthropic claimed to have uncovered millions of exchanges between its models and users identified at these companies through IP addresses and other metadata. These queries were characterized as "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use." While Anthropic has not yet responded to TechCrunch’s specific queries regarding Fable distillation, the pattern of accusations indicates a long-standing concern within the company about intellectual property protection.
However, the practice of "distillation" or using outputs from one model to train another is not unique to Chinese firms and exists in a gray area within the broader AI industry. Elon Musk, for instance, testified earlier this year that his company, xAI (formerly SpaceXAI), distilled OpenAI models to develop Grok, openly acknowledging that this practice was common within the industry. The line between legitimate research, developing synthetic datasets, and outright intellectual property theft through distillation remains ambiguously defined. Many researchers view the outputs of a model, once publicly accessible, as data that can be used for further training, similar to how human-generated content on the internet is used. The ethical and legal boundaries are still being negotiated, particularly concerning "open-weight" models which make their parameters publicly available, theoretically inviting further development and scrutiny.
The Crucial Role of Advanced Chips and Export Controls
Beyond the debate over distillation, the second core allegation—Moonshot’s alleged acquisition of banned advanced Nvidia chips—adds a critical dimension to the controversy. Kratsios specifically mentioned Grace Blackwell 300s (GB300s) and access to GB300-equipped servers in Thailand. These chips represent the pinnacle of AI computing power, and their export to China is strictly prohibited under U.S. regulations designed to curb China’s military and technological advancements.
The existence of a black market for these highly sought-after chips is well-documented. Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology, noted that despite export bans, illicit channels facilitate the flow of such technology. A notable instance of this was the indictment in May of the founder of Supermicro, a prominent U.S. server builder, for allegedly smuggling advanced chips into China. This incident underscores the significant challenges in enforcing export controls and preventing the diversion of critical technology.
Bresnick advocates for stronger "know your customer" (KYC) laws for data centers globally, arguing that "If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing." In 2024, President Joe Biden’s Department of Commerce proposed federal KYC rules for data centers, but progress on these regulations under the current administration remains unclear. Existing regulations already mandate that exporters shipping advanced chips abroad must "ensure they are only used for approved purposes," placing a burden of responsibility on U.S. companies to track and verify the end-use of their products. The alleged use of servers in Thailand further complicates enforcement, as it introduces a third-party jurisdiction where oversight might be weaker.
Broader Implications and the Future of AI Development
The allegations against Moonshot encapsulate the multifaceted challenges of the US-China technological competition. On one hand, they highlight the U.S.’s determination to protect its intellectual property and maintain its lead in frontier AI research. On the other, they underscore the difficulties in enforcing complex export controls and the blurred lines of ethical conduct in a rapidly evolving industry.
Braden Hancock’s perspective offers a nuanced counterpoint to the narrative of wholesale intellectual property theft: "[I]n general, Americans are understating the technical expertise of these Chinese teams. One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here." This view suggests that while some level of "distillation" or inspiration might occur, Chinese AI firms possess significant indigenous talent and capabilities that allow them to innovate independently, rather than merely copying Western advancements. The rapid emergence of highly capable Chinese models like Kimi K3, even if partially informed by existing models, points to a strong underlying research and development ecosystem.
The outcome of these allegations could have profound implications for the global AI ecosystem. If proven, they could lead to more stringent regulations on open-weight models, potentially stifling collaborative research and development. They could also escalate trade tensions, prompting further restrictions on technology transfer and increased scrutiny of international data center operations. The U.S. government’s stance reflects a growing concern that uncontrolled access to advanced AI capabilities by strategic rivals could pose national security risks, particularly in areas like military applications and surveillance technologies.
Ultimately, the controversy surrounding Moonshot’s Kimi K3 illuminates the urgent need for clear international norms and regulations concerning AI development, intellectual property, and responsible technology transfer. As AI continues to reshape industries and societies, the debate over who controls its foundational technologies and how those technologies are developed will remain at the forefront of geopolitical discourse. The current situation demands a careful balance between fostering innovation, protecting national interests, and promoting an ethical and transparent global AI environment. The lack of detailed evidence from the U.S. government and Moonshot’s silence only add to the complexity, leaving the AI community to grapple with serious accusations and their potential ramifications.







