arXiv:2607.21721stat.MLcs.LG2026-07

用历史重建结果训练先验会隐藏过自信,导致可信区间失效。

Priors learned from legacy reconstructions inherit undetectable overconfidence

论文配图:Priors learned from legacy reconstructions inherit undetectable overconfidence
图 1 · 摘自论文原文
  • 用旧方法的重建结果做先验,隐含错误假设却无法检测。
  • 在无法分辨的方向上,可信区间覆盖率远低于应有水平。
  • 需额外真实样本验证,适合地震、地下水等弱信号领域研究者。

当真实值稀缺(如地震与医学成像)时,常基于历史重建结果训练反问题的先验,并将不确定性视为数据驱动。理论上,若拥有无限多后验样本,其构成的档案即为正则化器;在可解析方向上,先验可优化;在盲区方向上,更新恒为恒等操作,错误假设永久留存。仅保留单个最佳重建的档案会抹除该区域的不确定性,使误差转化为过自信。这种假设随档案进入,以报告方差形式输出,部署中无法验证。两个仅在盲区不同的真值共享相同观测数据,任何仅依赖调查与档案的程序都无法同时保证有限盲区且覆盖所有不可区分的真值。我们提出可分辨性判据,明确受影响方向,说明测试所需参考真值数量,并据此构建声称包含真实值的区间——即使先验错误也成立。合成实验显示,地震与地下水场景下,档案训练先验在盲区方向的覆盖率显著低于可解析方向,而真值训练的对照组无此偏差。地震中随机子空间产生类似结果,说明差异由算子决定;地下水场景中,利用调查与少量钻孔重构建档,使先验报告向真实值靠拢,但仅限于钻孔可达方向。

原文摘要 · Abstract (English)

Where truths are scarce (e.g., seismic and medical imaging), a prior for an ill-posed inverse problem is trained on an archive of legacy reconstructions---an older method's outputs---and its uncertainty is treated as data-driven. In the population limit, an archive of posterior samples is the regularizer that produced it, advanced one expectation-maximization step toward the truth. On directions the operator resolves, it improves the assumption; on its blind subspace, the step is the identity, so the assumption survives unchanged however often it is rebuilt. An archive of single-best reconstructions, one per survey, keeps no spread there: the blind interval collapses whatever the penalty was, so error becomes overconfidence. The assumption enters as the archive and leaves as a reported spread, and nothing in deployment tests it. Two truths differing only there share the data law, and no procedure using survey and archive alone can both report a finite blind interval and guarantee coverage over indistinguishable truths. The question requires information the survey does not carry. We provide a resolvability statement that names affected directions from the operator, state how many reference truths are needed to test a prior on them, and use those references to build an interval that contains the truth as claimed, even if the prior is wrong. On synthetic experiments with seismic and groundwater operators, the archive-trained prior's intervals contain the truth less often on the blind subspace than on resolved directions, while the truth-trained control shows no gap of that sign or size. On seismic, a random subspace of the same size gives the same result, so the separation follows the operator. On groundwater, rebuilding the archive using the survey and a handful of boreholes brings the prior's reports toward the truth on directions those boreholes reach, while leaving the rest unchanged.

反问题先验建模可信区间地震成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。