arXiv:2508.17761cs.LGstat.ML2025-08被引 3

评估回归模型不确定性校准质量,发现现有指标常矛盾,推荐用ENCE和CWC。

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

  • 系统梳理并独立测试多种校准指标,排除模型方法干扰。
  • 实测显示多数指标对同一结果评价冲突,部分甚至相反。
  • 建议研究者优先选用ENCE和CWC,避免结果误导。

在安全关键应用中,数据驱动模型不仅需准确,还需提供可靠的不确定性估计。这一特性被称为校准,对风险感知决策至关重要。尽管回归任务已涌现大量校准度量与校准方法,但这些度量在定义、假设和量纲上差异显著,导致跨研究结果难以解释与比较。此外,多数校准方法仅用少量度量评估,无法判断改进是否在不同校准概念下通用。本文从文献中系统提取并分类回归校准度量,并独立于具体建模方法或校准策略进行基准测试。通过真实世界、合成及人为校准偏差数据的受控实验,我们发现校准度量常产生冲突结果。分析表明:许多度量对同一校准结果评价不一致,某些甚至得出相反结论。这种不一致性令人担忧,可能允许选择性使用度量制造虚假成功印象。我们识别出期望归一化校准误差(ENCE)和覆盖率宽度准则(CWC)在测试中表现最可靠。研究强调了度量选择在校准研究中的关键作用。

原文摘要 · Abstract (English)

In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonly referred to as calibration, is essential for risk-aware decision-making. In regression a wide variety of calibration metrics and recalibration methods have emerged. However, these metrics differ significantly in their definitions, assumptions and scales, making it difficult to interpret and compare results across studies. Moreover, most recalibration methods have been evaluated using only a small subset of metrics, leaving it unclear whether improvements generalize across different notions of calibration. In this work, we systematically extract and categorize regression calibration metrics from the literature and benchmark these metrics independently of specific modelling methods or recalibration approaches. Through controlled experiments with real-world, synthetic and artificially miscalibrated data, we demonstrate that calibration metrics frequently produce conflicting results. Our analysis reveals substantial inconsistencies: many metrics disagree in their evaluation of the same recalibration result, and some even indicate contradictory conclusions. This inconsistency is particularly concerning as it potentially allows cherry-picking of metrics to create misleading impressions of success. We identify the Expected Normalized Calibration Error (ENCE) and the Coverage Width-based Criterion (CWC) as the most dependable metrics in our tests. Our findings highlight the critical role of metric selection in calibration research.

不确定性校准回归度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。