source: kdnuggets: humanity’s last exam is a distraction

level: technical

humanity's last exam is a benchmark built by the center for ai safety and scale ai to test advanced ai systems. it contains over 2,500 expert-level questions across more than a hundred academic fields like physics, math, and humanities. the questions require complex reasoning and deep understanding, not just memorization or simple retrieval. even top models like gpt, gemini, and claude score only around 45-50% accuracy, often showing overconfidence in wrong answers.

opinions on hle are split into three main groups. about 60% of experts see it as useful and necessary because older benchmarks like mmlu became saturated, with models scoring over 90%. they value that hle tests whether ai admits uncertainty instead of hallucinating. around 30% call it a distraction, arguing it focuses on obscure academic knowledge rather than real-world ai performance. a smaller group claims the benchmark has errors in some answers, especially in niche chemistry and math questions, which advanced ai systems themselves have detected.

despite the debate, most experts do not dismiss hle entirely. they criticize its dramatic name as marketing hype but acknowledge its role in comparing model reasoning and memory. the benchmark is not seen as a path to artificial general intelligence, but as an ambitious tool to identify which ai models have the strongest logical capabilities.

why it matters: it shows how ai evaluation is evolving beyond saturated tests, helping data scientists understand model limitations in reasoning and uncertainty handling.


source: kdnuggets: humanity’s last exam is a distraction