提出统一风险视角,精准拆分预测不确定性并评估其可靠性。
A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

- 以后验风险定义不确定性,融合函数不确定与优化偏差
- 基于半合成数据直接计算真实贝叶斯与随机不确定性
- 揭示模型表现好但不确定性不准确的常见问题,适合安全场景研究者
可靠不确定性估计对安全敏感应用至关重要,需区分认知不确定性与随机不确定性,但现有文献定义不一,难以评估方法准确性。由于真实认知不确定性通常不可知,现有评估多依赖分布外检测等代理任务,无法提供完整真值且难以洞察不确定性结构。本文提出将不确定性统一定义为点态后验风险——在给定数据下,合理真实函数分布上预测器的期望损失。该视角结合贝叶斯函数不确定性与估计器偏离后验均值的偏差,涵盖模型误设与优化误差。基于此构建理论支持的基准,利用具有真实协变量和已知生成过程的半合成数据,可直接计算理想认知与随机不确定性。避免代理评估,实现对不确定性估计的细粒度分析。实验发现:高预测精度并不保证不确定性拆分可靠。基准揭示了不同方法间实际差异,识别出与理想目标对齐的方法,并暴露其对数据集和建模选择的敏感性。
原文摘要 · Abstract (English)
Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric uncertainty, yet these uncertainty types are not defined consistently across the literature, making it difficult to assess whether a method produces accurate uncertainty estimates. Evaluation is further complicated by the fact that ground-truth epistemic uncertainty is typically unavailable. Existing benchmarks therefore mostly rely on proxy tasks such as out-of-distribution detection, which do not provide complete ground-truth uncertainty targets and offer limited insight into the structure and quality of uncertainty estimates. We propose a unified definition of uncertainty as pointwise posterior risk, the expected loss of a predictor under the distribution of plausible ground-truth functions given the data. This view combines Bayesian uncertainty over functions with estimator-dependent deviations from the posterior mean, capturing effects such as misspecification and optimization error. This formulation constitutes the foundation of a theory-backed benchmark that enables direct computation of oracle epistemic and aleatoric uncertainty using semi-synthetic datasets with real covariates and known generative processes. By avoiding proxy evaluations, the benchmark enables fine-grained analysis of uncertainty estimates. Empirically, we find that accurate prediction does not guarantee reliable uncertainty disentanglement. The benchmark reveals practically useful differences between methods, identifying approaches with meaningful alignment to oracle uncertainty targets while exposing sensitivity to datasets and modeling choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。