提出新方法同时校准数据噪声和模型不确定,提升预测区间可靠性。
CLEAR: Calibrated Learning for Epistemic and Aleatoric Risk
- 用两个参数分别控制两类不确定性,统一校准预测区间
- 在17个真实数据集上,区间宽度平均缩小28.3%和17.5%
- 特别适合高噪声或数据稀缺场景,兼容多种模型框架
准确的不确定性量化对可靠预测建模至关重要。现有方法通常只处理测量噪声带来的似然不确定性或数据不足导致的认知不确定性,无法兼顾二者。本文提出CLEAR,一种包含两个独立参数 $γ_1$ 和 $γ_2$ 的校准方法,用于联合建模两类不确定性,并提升回归任务中预测区间的条件覆盖率。CLEAR可与任意一对似然与认知不确定性估计器兼容;我们展示了其如何与(i)分位数回归(用于似然不确定性)和(ii)来自可预测性-可计算性-稳定性(PCS)框架的集成模型(用于认知不确定性)结合使用。在17个多样化的真实世界数据集上,相较于两个独立校准基线,CLEAR平均将区间宽度分别缩小28.3%和17.5%,同时保持名义覆盖水平。当应用于深度集成(认知)与同步分位数回归(似然)时也获得类似改进。在高似然或高认知不确定性场景下,优势尤为显著。
原文摘要 · Abstract (English)
Accurate uncertainty quantification is critical for reliable predictive modeling. Existing methods typically address either aleatoric uncertainty due to measurement noise or epistemic uncertainty resulting from limited data, but not both in a balanced manner. We propose CLEAR, a calibration method with two distinct parameters, $γ_1$ and $γ_2$, to combine the two uncertainty components and improve the conditional coverage of predictive intervals for regression tasks. CLEAR is compatible with any pair of aleatoric and epistemic estimators; we show how it can be used with (i) quantile regression for aleatoric uncertainty and (ii) ensembles drawn from the Predictability-Computability-Stability (PCS) framework for epistemic uncertainty. Across 17 diverse real-world datasets, CLEAR achieves an average improvement of 28.3\% and 17.5\% in the interval width compared to the two individually calibrated baselines while maintaining nominal coverage. Similar improvements are observed when applying CLEAR to Deep Ensembles (epistemic) and Simultaneous Quantile Regression (aleatoric). The benefits are especially evident in scenarios dominated by high aleatoric or epistemic uncertainty. Project page: https://unco3892.github.io/clear/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。