source: techcrunch ai: experts say exploiting anthropic’s fable isn’t how kimi k3 got so good
level: technical
white house science advisor michael kratsios claimed that chinese company moonshot built its kimi k3 model by copying anthropic's fable llm using restricted chips. kratsios called it large-scale covert industrial distillation and unacceptable. treasury secretary scott bessent also mentioned finding watermarks of us models on chinese ones, though details remain unclear. moonshot did not respond to questions, and kratsios provided no further evidence.
experts are skeptical that distillation alone explains kimi k3's capabilities. braden hancock of the laude institute noted the timeline is too short, as fable was only public since july 1. nathan lambert of the allen institute for ai argued that distillation's impact is decreasing as chinese models near the frontier and training shifts to reinforcement learning. he said supervised fine-tuning, which teaches models manners, is less crucial for complex capabilities.
distillation involves querying a model to generate training data, but replicating fable-like performance would likely need expensive reinforcement learning with millions of agents, making it impractical via api. anthropic previously accused moonshot of systematic distillation, but the practice is common industry-wide. hancock emphasized the technical expertise of chinese teams, noting moonshot's founder studied at carnegie mellon. the debate also involves alleged use of banned nvidia chips, with calls for stricter know-your-customer rules for data centers.
why it matters: understanding how leading models are built affects ai policy, export controls, and the global competition in developing advanced ai systems.
source: techcrunch ai: experts say exploiting anthropic’s fable isn’t how kimi k3 got so good