source: arXiv Statistics ML: Prior laundering: learned priors with inherited, undetectable overconfidence

level: research

Learned generative priors are widely used for Bayesian inverse problems in fields like seismic and medical imaging. When ground truth data is scarce, practitioners often train on archives of legacy reconstructions, a process the authors call prior laundering. The resulting prior appears data-driven but actually inherits beliefs from the old reconstructions. This creates a hidden risk: the prior’s uncertainty can be overconfident in directions where the original measurements were uninformative.

The paper proves that this overconfidence is undetectable during deployment. When measurements lack information, the posterior reverts to the prior, so reported uncertainty reflects the archive, not the data. Standard self-consistency checks like simulation-based calibration pass regardless of the prior’s accuracy. The inherited belief is exactly the old regularizer advanced by one expectation–maximization step—improved where data resolve, frozen where they cannot, leading to tighter uncertainty than the truth warrants.

This finding matters for any application where learned priors are trained on reconstructed rather than true images. It shows that even rigorous calibration cannot reveal when a prior is too narrow. The work connects to broader concerns about feedback loops in machine learning, where models trained on synthetic or processed data can amplify hidden biases. Researchers should verify whether their training data contains genuine ground truth or merely legacy estimates.

why it matters: It reveals a blind spot in Bayesian deep learning: priors trained on reconstructions can be overconfident in ways that standard diagnostics miss, risking misleading uncertainty in critical imaging tasks.


source: arXiv Statistics ML: Prior laundering: learned priors with inherited, undetectable overconfidence