source: Google DeepMind: Piloting the world's first double-blind AI evaluations
level: technical
google deepmind ran the first double-blind evaluation of a proprietary frontier ai model. the test used a gemini flash lite model and confidential benchmarks inside a cryptographic box. partners included singapore ai safety institute, openmined, averi, and mlcommons. the setup prevents the model from seeing test questions ahead of time. it also stops evaluators from seeing model weights. this addresses benchmark contamination, where models score higher because they already saw the questions.
the method uses confidential space in google cloud's confidential computing portfolio. cryptographic verification shows both the evaluation data and the model stay private to their owners. google cannot see the test prompts, and evaluators cannot see the model weights. this removes the old tradeoff between handing over prompts or model weights. the pilot tested a gemini flash lite model against confidential benchmarks. no specific scores or contamination rates were published in the blog post.
benchmark contamination is a known problem in ai evaluation. if a model trains on test data, its scores become unreliable. this matters for safety tests, especially in cybersecurity or government use. double-blind evaluation could let independent groups test models without risking data sovereignty. the approach builds on existing zero-logging protocols and contracts. it adds technical and cryptographic safeguards. the pilot is a step toward more trustworthy ai benchmarks for policymakers and researchers.
why it matters: double-blind evaluation prevents ai models from memorizing test answers, so benchmark scores reflect true capabilities and safety.
source: Google DeepMind: Piloting the world's first double-blind AI evaluations