arXiv:2510.21273stat.MLcs.LG2025-10

提出新方法提升多输出概率回归校准度,避免决策偏差。

Enforcing Calibration in Multi-Output Probabilistic Regression with Pre-rank Regularization

  • 用预排序正则化惩罚投影后概率积分变换偏离均匀分布
  • 18个真实数据集上显著改善校准性且不降低预测精度
  • 支持任意预排序函数,含主成分分析新方法

概率模型需良好校准以支持可靠决策。尽管单输出回归的校准已研究充分,多输出回归中的多变量校准仍具挑战。现有文献主要聚焦基于预排序函数的诊断工具,这些函数将多变量预测-观测对投影为一维摘要,用于检测特定类型的校准偏差。本文超越诊断,提出一种通用正则化框架,在训练中强制实现多变量校准,适用于任意预排序函数。该框架涵盖最高密度区域校准和耦合校准等方法。通过惩罚投影后的概率积分变换(PIT)偏离均匀分布来实现校准,并可作为正则项加入任意概率预测器的损失函数。我们提出一种联合强制边际与多变量预排序校准的正则化损失。还引入一种基于主成分分析(PCA)的预排序,捕捉预测分布方差最大方向上的校准特性,同时实现降维。在18个真实世界多输出回归数据集上,未正则化模型均存在系统性校准偏差,而我们的方法在所有预排序函数下均显著提升校准性,且不牺牲预测准确性。

原文摘要 · Abstract (English)

Probabilistic models must be well calibrated to support reliable decision-making. While calibration in single-output regression is well studied, defining and achieving multivariate calibration in multi-output regression remains considerably more challenging. The existing literature on multivariate calibration primarily focuses on diagnostic tools based on pre-rank functions, which are projections that reduce multivariate prediction-observation pairs to univariate summaries to detect specific types of miscalibration. In this work, we go beyond diagnostics and introduce a general regularization framework to enforce multivariate calibration during training for arbitrary pre-rank functions. This framework encompasses existing approaches such as highest density region calibration and copula calibration. Our method enforces calibration by penalizing deviations of the projected probability integral transforms (PITs) from the uniform distribution, and can be added as a regularization term to the loss function of any probabilistic predictor. Specifically, we propose a regularization loss that jointly enforces both marginal and multivariate pre-rank calibration. We also introduce a new PCA-based pre-rank that captures calibration along directions of maximal variance in the predictive distribution, while also enabling dimensionality reduction. Across 18 real-world multi-output regression datasets, we show that unregularized models are consistently miscalibrated, and that our methods significantly improve calibration across all pre-rank functions without sacrificing predictive accuracy.

概率回归校准多输出正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。