arXiv:2505.08784stat.MLcs.LG2025-05被引 11

提出PCS-UQ框架,实现更可靠且自适应的不确定性量化。

PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework

  • 基于可预测性、可计算性与稳定性原则,筛选模型并结合自助采样捕捉变异性。
  • 在17个真实回归数据集上覆盖率达目标值,区间宽度优于或等同于最优对比方法。
  • 适用于高风险场景,尤其适合需要稳定子组覆盖的机器学习应用。

随着机器学习进入高风险领域,可信的不确定性量化(UQ)对安全至关重要。本文提出基于可预测性、可计算性与稳定性(PCS)原则的PCS-UQ框架,用于真实数据科学。从候选模型集合出发,通过严格的预测检验筛选不合适的模型,并利用自助样本捕获样本间变异性和算法不稳定性。引入一种新型乘法校准方案以增强局部自适应性,可视为置信推断中的新评分机制。构建了包含17个真实世界回归数据集的基准,其手动划分的子组中,PCS-UQ保持目标覆盖率,同时区间宽度优于或等同于使用最优算法的置信方法;在子组覆盖上表现更优。值得注意的是,该方法在区间宽度和子组一致性上均具竞争力。在6个分类数据集上,预测集大小平均减少20%。为适配深度学习,提出计算高效的变体,避免昂贵重训练,在3个计算机视觉基准上,预测集大小较置信基线减少20%。最后,理论证明修改后的PCS-UQ算法在交换性假设下仍保持有效覆盖率,属于分裂置信推断的一种形式。

原文摘要 · Abstract (English)

As machine learning (ML) enters high-stakes domains, trustworthy uncertainty quantification (UQ) is essential for safety. In this paper we introduce PCS-UQ, a framework based on the Predictability, Computability, and Stability (PCS) principles for veridical data science. Starting with a candidate set of models or algorithms, PCS-UQ integrates a rigorous prediction-check to screen out unsuitable models in the set and utilizes bootstrap samples in order to capture both inter-sample variability and algorithmic instability for the prediction-checked algorithms. We then introduce a novel multiplicative calibration scheme to enhance local adaptivity, which can be viewed as a new score in conformal prediction. Moreover, we produce a compilation of 17 real-world regression datasets with manually constructed subgroups. On this benchmark, PCS-UQ maintains the target coverage while outperforming or matching conformal methods equipped with oracle-selected algorithms in interval width. PCS-UQ achieves consistent subgroup coverage, outperforming these oracle-selected conformal methods. Notably, PCS-UQ stands out in achieving both competitive interval widths and consistent subgroup coverage. Across 6 classification datasets, PCS-UQ reduces prediction set sizes by 20\%. To scale the framework for deep learning, we propose computationally efficient variants that bypass expensive retraining. On three computer vision benchmarks, these variants reduce prediction set sizes by 20\% over conformal baselines. Finally, we provide a theoretical proof that a modified PCS-UQ algorithm preserves valid coverage under exchangeability as a form of split conformal inference.

不确定性量化置信推断机器学习安全模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。