source: arxiv machine learning: auditing the audit: five failure modes in benchmark-validity audits

level: research

governance frameworks often require ai providers and auditors to submit perturbation-based construct-validity audits as proof of model safety. these audits are meant to check if a benchmark really measures what it claims. but the audit process itself can break down in ways that are invisible in the final reported numbers. the authors identify five failure modes that can silently manufacture false conclusions.

the five failure modes cover issues like using the wrong perturbation, misinterpreting statistical results, or ignoring how model training data overlaps with benchmarks. the researchers tested these failures on two open-weight instruction-tuned models across five safety benchmarks. they applied a six-point due-diligence gate to every result. not a single cell passed as confirmatory. every result fell into a non-confirmatory bucket, meaning the audit could not be trusted.

the proposed due-diligence gate is a checklist for withholding or disclosing audit evidence. it is not a replacement for standard audit methods but a supplement to raise the bar for assurance-grade evidence. the study is limited to a small case study and the failure taxonomy is not exhaustive. still, it shows that current audit practices can easily miss critical flaws, and that stronger protocols are needed to make benchmark validity checks reliable.

why it matters: if audit evidence is silently flawed, ai safety claims based on benchmarks become unreliable, risking deployment of unsafe models.


source: arxiv machine learning: auditing the audit: five failure modes in benchmark-validity audits