arXiv:2502.16299cs.LGcs.AI2025-02被引 10

提出新方法检验不确定性集合是否可靠

A calibration test for evaluating set-based epistemic uncertainty representations

  • 用可变组合方式测试预测集是否校准
  • 在合成与真实数据上验证了方法有效性
  • 适合关注模型不确定性的研究者

准确表示认知不确定性是机器学习中的关键挑战。常用方法是使用凸概率预测子集(即可信集),通过集成或专门监督学习构建,其不确定性可通过集合大小或成员间分歧衡量。理论上,这些集合应包含真实数据生成分布。为此,我们采用最强校准概念作为必要条件,提出一种新型统计检验方法:判断集合内预测的凸组合是否在分布上校准。与以往方法不同,该框架允许组合随输入实例变化,承认不同成员在不同输入区域校准程度不同。通过合理评分规则学习该组合,并基于可微分、基于核的校准误差估计器,构建非参数检验流程。在合成与真实数据实验中均验证了捕捉实例级差异的优势。

原文摘要 · Abstract (English)

The accurate representation of epistemic uncertainty is a challenging yet essential task in machine learning. A widely used representation corresponds to convex sets of probabilistic predictors, also known as credal sets. One popular way of constructing these credal sets is via ensembling or specialized supervised learning methods, where the epistemic uncertainty can be quantified through measures such as the set size or the disagreement among members. In principle, these sets should contain the true data-generating distribution. As a necessary condition for this validity, we adopt the strongest notion of calibration as a proxy. Concretely, we propose a novel statistical test to determine whether there is a convex combination of the set's predictions that is calibrated in distribution. In contrast to previous methods, our framework allows the convex combination to be instance dependent, recognizing that different ensemble members may be better calibrated in different regions of the input space. Moreover, we learn this combination via proper scoring rules, which inherently optimize for calibration. Building on differentiable, kernel-based estimators of calibration errors, we introduce a nonparametric testing procedure and demonstrate the benefits of capturing instance-level variability on of synthetic and real-world experiments.

不确定性建模校准测试可信集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。