arXiv:2508.18001cs.LGstat.ML2025-08被引 1

用合理评分统一量化模型不确定性,通用性强且效果优。

A Novel Framework for Uncertainty Quantification via Proper Scores for Classification and Beyond

  • 基于严格合理评分构建偏差-方差分解,统一处理各类任务不确定性。
  • 在图像、音频、语言生成中,对大模型不确定性估计超越现有方法。
  • 可推广至校准性评估,适合追求模型可信度的研究者使用。

本博士论文提出一种基于合理评分的机器学习不确定性量化新框架。不确定性量化是实现可信可靠机器学习应用的基础。传统方法多针对特定任务,难以迁移。合理评分是使目标分布预测最优的损失函数,适用于回归、分类及生成建模等任务。本文通过泛函Bregman散度建立严格合理评分的通用偏差-方差分解,揭示了认知不确定性、随机不确定性与模型校准之间的理论联系。研究采用核评分(kernel score)评估图像、音频、自然语言生成等领域的样本生成模型,提出一种新颖的大语言模型不确定性估计方法,性能优于现有基准。进一步将校准-锐度分解推广至分类以外的任务,定义合理校准误差,并提出分类任务中新校准误差估计器及基于风险的平方校准误差估计器比较方法。最后,对核球面评分进行分解,实现生成图像模型更精细、可解释的评估。

原文摘要 · Abstract (English)

In this PhD thesis, we propose a novel framework for uncertainty quantification in machine learning, which is based on proper scores. Uncertainty quantification is an important cornerstone for trustworthy and reliable machine learning applications in practice. Usually, approaches to uncertainty quantification are problem-specific, and solutions and insights cannot be readily transferred from one task to another. Proper scores are loss functions minimized by predicting the target distribution. Due to their very general definition, proper scores apply to regression, classification, or even generative modeling tasks. We contribute several theoretical results, that connect epistemic uncertainty, aleatoric uncertainty, and model calibration with proper scores, resulting in a general and widely applicable framework. We achieve this by introducing a general bias-variance decomposition for strictly proper scores via functional Bregman divergences. Specifically, we use the kernel score, a kernel-based proper score, for evaluating sample-based generative models in various domains, like image, audio, and natural language generation. This includes a novel approach for uncertainty estimation of large language models, which outperforms state-of-the-art baselines. Further, we generalize the calibration-sharpness decomposition beyond classification, which motivates the definition of proper calibration errors. We then introduce a novel estimator for proper calibration errors in classification, and a novel risk-based approach to compare different estimators for squared calibration errors. Last, we offer a decomposition of the kernel spherical score, another kernel-based proper score, allowing a more fine-grained and interpretable evaluation of generative image models.

不确定性量化合理评分生成模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。