不同任务需用不同不确定性度量方法
Uncertainty Quantification for Machine Learning: One Size Does Not Fit All
- 用可调损失函数构建灵活的不确定性度量框架
- 选择性预测应匹配任务损失,OoD检测互信息最优
- 主动学习中零一损失的后验不确定性更优
在安全关键应用中,准确量化预测不确定性至关重要。现有多种不确定性度量方法常宣称优于其他方法。本文认为不存在通用最优度量,应根据具体任务定制。为此,我们提出一个灵活的不确定性度量族,可区分二阶分布下的总不确定性、随机性不确定性和认知不确定性。这些度量可通过特定损失函数(即合理评分规则)实例化,以控制其特性。实验表明,对于选择性预测任务,评分规则应与任务损失一致;在分布外检测任务中,广泛使用的认知不确定性度量——互信息表现最佳;在主动学习场景下,基于零一损失的认知不确定性始终优于其他度量。
原文摘要 · Abstract (English)
Proper quantification of predictive uncertainty is essential for the use of machine learning in safety-critical applications. Various uncertainty measures have been proposed for this purpose, typically claiming superiority over other measures. In this paper, we argue that there is no single best measure. Instead, uncertainty quantification should be tailored to the specific application. To this end, we use a flexible family of uncertainty measures that distinguishes between total, aleatoric, and epistemic uncertainty of second-order distributions. These measures can be instantiated with specific loss functions, so-called proper scoring rules, to control their characteristics, and we show that different characteristics are useful for different tasks. In particular, we show that, for the task of selective prediction, the scoring rule should ideally match the task loss. On the other hand, for out-of-distribution detection, our results confirm that mutual information, a widely used measure of epistemic uncertainty, performs best. Furthermore, in an active learning setting, epistemic uncertainty based on zero-one loss is shown to consistently outperform other uncertainty measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。