用可减少的错误评估不确定性,更准确判断模型是否该拒答。
Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

- 基于可减少误差设计新评估框架,理论证明最优拒答策略是两类不确定性的加权阈值。
- 发现现有相关性指标无法预测不确定性组件的实际效果,可能误导模型改进方向。
- 建议用风险-后悔-覆盖率三维表面诊断不确定性分解效果,适合关注决策可靠性的研究者。
当前对认知不确定性(epistemic uncertainty)的评估主要依赖于分布外检测和主动学习等任务,但这些任务的贝叶斯最优决策策略与常用不确定性评分并不一致。本文基于认知拒答框架,提出以可减少误差(regret)作为评估标准。将选择性预测建模为覆盖度、期望风险与后悔之间的约束优化问题,理论证明最优选择器是真实本征不确定性与认知不确定性凸组合的阈值形式。这一理论统一揭示了近期不确定性解耦研究的弱点:我们证明,学习组件间的标准相关性指标未必反映其实际操作效用。因此,我们提出应通过评估可实现风险、后悔与覆盖度的联合表面来诊断解耦效果与实用性。在具有密集人工标注的数据集上对主流方法进行基准测试显示,基于决策论的排序与代理任务排序存在显著差异,包括方法间排名完全颠倒的情况,即某些方法在一个指标中位列榜首,在另一指标中却垫底。
原文摘要 · Abstract (English)
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。