小样本下贝叶斯深度学习方法排名不可靠,需用不确定性评估真伪。
Unstable Rankings in Bayesian Deep Learning Evaluation

- 用分层贝叶斯模型把评估指标当随机变量处理,捕捉数据波动影响。
- 同一方法在不同数据集上排名差异大:n=50时MCD劣于Ensemble概率达1.0,另一数据集n=500仍低于0.95。
- 提出可检测差异曲线,帮助判断当前数据量能否可靠比较方法优劣。
标准贝叶斯深度学习评估假设指标估计可靠,但我们在数据稀缺时发现该假设失效。方法排名不仅在小样本(n)下不可靠,且受数据集影响显著,点估计无法揭示这种差异:同一对比在某数据集上n=50时MCD劣于Ensemble的概率为1.000,而在另一数据集上n=500时仍低于0.95。跨五个回归数据集分析表明,不存在普适的样本量阈值,因此必须进行数据集特定的后验推断。为此,我们采用具有方法特异性方差的贝叶斯分层模型,将评估指标视为数据实现下的随机变量,并引入预测最小可检测差异曲线,评估给定训练规模下观测差距是否可检出。在六种贝叶斯深度学习方法和五组数据上的结果表明,在低数据环境下,不确定性感知评估至关重要——当前证据显示的方法优势与预测可检测性可能严重偏离。本框架为从业者提供依据,判断其评估数据是否足够再得出方法优劣结论。
原文摘要 · Abstract (English)
Standard evaluations of Bayesian deep learning methods assume that metric estimates are reliable, but we show this assumption fails under data scarcity. Method rankings are not only unreliable at small $n$, but also dataset-dependent in ways that point estimates cannot reveal: the same method comparison yields $P(\mathrm{MCD} \prec \mathrm{Ensemble}) = 1.000$ at $n = 50$ on one dataset and remains below $0.95$ even at $n = 500$ on another. Across the datasets we consider, no universal sample size threshold exists, which is precisely why dataset-specific posterior inference is necessary. To address this, we use a Bayesian hierarchical model with method-specific variances to treat evaluation metrics as random variables across data realizations, and we use a predictive Minimum Detectable Difference curve to assess whether an observed gap would be detectable at a given training size. Across six Bayesian deep learning methods and five regression datasets, our results show that uncertainty-aware evaluation is necessary in low-data settings, because current evidence for method superiority and predictive detectability at the same training size can diverge substantially. Our framework provides practitioners with principled tools to determine whether their evaluation data is sufficient before drawing conclusions about method superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。