arXiv:2606.26990cs.LGcs.AI2026-06

提出决策对齐评估方法,让不确定性量化更贴近实际应用价值

Decision-Aligned Evaluation of Uncertainty Quantification

  • 用决策对齐标准检验评估指标与实际决策的匹配度
  • 发现主流评估指标常与真实决策需求脱节
  • 新提出的加权效用指标能准确反映决策实际收益

机器学习中的不确定性估计通常采用负对数似然、期望校准误差等通用指标进行评估,但这些指标表现良好并不意味着在下游决策中真正有用。本文提出决策对齐准则,揭示哪些评估指标能真正反映下游决策价值。基于该框架,我们发现许多常用不确定性指标要么与常见决策问题不匹配,要么隐含病态的先验假设。为此,我们提出一种特殊的严格评分规则——先验加权效用指标,实现决策对齐的不确定性评估。在基准实验和真实案例研究中,该方法始终与实际决策效用保持一致,而传统指标则否。研究揭示了当前不确定性量化评估协议的缺陷,并为构建面向决策的评估体系提供了理论扩展。

原文摘要 · Abstract (English)

Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common decision problems or encode pathological prior beliefs about the downstream task. We then propose prior-weighted utility metrics, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface flaws in the current UQ evaluation protocol and offer a principled extension of existing metrics toward decision-relevant UQ evaluation.

不确定性量化评估方法决策对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。