arXiv:2603.08495cs.LGstat.ML2026-03被引 1

提出高效可信区间预测方法,让大模型也能快速估算不确定性。

Efficient Credal Prediction through Decalibration

  • 基于相对似然和去校准技术,为每类输出概率区间。
  • 在多种任务中表现优异,覆盖效率与异常检测均超现有方法。
  • 首次实现对TabPFN、CLIP等复杂模型的可信区间预测。

可靠地表示不确定性对现代机器学习在安全关键场景的应用至关重要。近年来,可信集(即概率分布的凸集)被提出作为表征认知不确定性的一种合适方法。然而,与其他认知不确定性方法类似,训练可信预测器计算复杂,通常需要(重新)训练模型集成,导致其难以应用于复杂模型如基础模型和多模态系统。为此,我们提出一种基于相对似然并受概率分类器校准技术启发的高效可信预测方法。对于每个类别标签,该方法预测一个可能的概率区间。通过提出称为‘去校准’的技术,生成这些区间的上下界。大量实验表明,该方法在多样化任务中表现出色,包括覆盖率-效率评估、分布外检测以及上下文学习。特别地,我们首次实现了对TabPFN和CLIP等架构的可信预测,而此前此类模型构建可信集尚不可行。

原文摘要 · Abstract (English)

A reliable representation of uncertainty is essential for the application of modern machine learning methods in safety-critical settings. In this regard, the use of credal sets (i.e., convex sets of probability distributions) has recently been proposed as a suitable approach to representing epistemic uncertainty. However, as with other approaches to epistemic uncertainty, training credal predictors is computationally complex and usually involves (re-)training an ensemble of models. The resulting computational complexity prevents their adoption for complex models such as foundation models and multi-modal systems. To address this problem, we propose an efficient method for credal prediction that is grounded in the notion of relative likelihood and inspired by techniques for the calibration of probabilistic classifiers. For each class label, our method predicts a range of plausible probabilities in the form of an interval. To produce the lower and upper bounds of these intervals, we propose a technique that we refer to as decalibration. Extensive experiments show that our method yields credal sets with strong performance across diverse tasks, including coverage-efficiency evaluation, out-of-distribution detection, and in-context learning. Notably, we demonstrate credal prediction on models such as TabPFN and CLIP -- architectures for which the construction of credal sets was previously infeasible.

不确定性估计可信集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。