source: Hugging Face Blog: What We Learned by Reproducing 2,200 papers from ICML
level: technical
A 19-day hackathon organized by Hugging Face and alphaXiv used coding agents to reproduce papers from ICML 2026. Over 1,200 participants published 6,816 logbooks covering 2,226 papers, about a third of the conference. The agents read papers, wrote code, ran experiments, and reported results. An automated judge then classified each claim as verified, falsified, toy, or inconclusive. The effort aimed to test whether AI agents could audit research at scale.
The results showed that 51% of examined papers had at least one verified claim, with 266 papers fully reproduced. However, 23% had at least one falsified or contested claim, including 49 papers where all claims were falsified. Notably, 242 papers had conflicting verdicts from independent teams. Some falsifications were confirmed by authors, such as a paging algorithm whose robustness term grew logarithmically instead of remaining constant. Other errors included using forward KL divergence instead of reverse KL and padding tokens inflating evaluation metrics.
The hackathon revealed that human oversight remains crucial. Agents often got stuck in loops or missed scale-dependent behavior, and some falsifications were themselves flawed due to arithmetic errors. The most reliable results came from human-steered workflows. The organizers argue that humans should manage agents like a principal investigator guides graduate students. All logbooks, verdicts, and traces are public, making this the largest claim-by-claim audit of a machine learning conference to date.
why it matters: This shows that AI agents can audit research at scale, but human judgment is still needed to catch errors and interpret results.
source: Hugging Face Blog: What We Learned by Reproducing 2,200 papers from ICML