arXiv:2601.22895cs.LG2026-01被引 1

用预排序正则化提升多变量预测的校准性,兼顾准确性

Calibrated Multivariate Distributional Regression with Pre-Rank Regularization

  • 基于预排序函数设计训练时正则化,实现多变量校准
  • 在18个真实数据集上显著改善校准性能,且不降低预测精度
  • 提出基于PCA的新型预排序,可发现传统方法遗漏的依赖结构错误

概率预测的目标是在保证校准性的前提下,提供尽可能信息丰富的预测分布。尽管单变量场景已有显著进展,多变量校准仍具挑战。近期工作引入预排序函数(pre-rank functions),作为多变量预测与观测的标量投影,用于灵活诊断校准中的特定问题,但其应用多限于事后评估。本文提出一种基于正则化的校准方法,在多变量分布回归模型训练过程中利用预排序函数强制实现多变量校准。我们进一步提出一种基于主成分分析(PCA)的新型预排序,将预测值投影到预测分布的主方向上。通过仿真研究和在18个真实世界多输出回归数据集上的实验表明,该方法显著提升了多变量预排序校准效果,且未牺牲预测准确性;同时,基于PCA的预排序能揭示现有预排序未能检测到的依赖结构误建模问题。

原文摘要 · Abstract (English)

The goal of probabilistic prediction is to issue predictive distributions that are as informative as possible, subject to being calibrated. Despite substantial progress in the univariate setting, achieving multivariate calibration remains challenging. Recent work has introduced pre-rank functions, scalar projections of multivariate forecasts and observations, as flexible diagnostics for assessing specific aspects of multivariate calibration, but their use has largely been limited to post-hoc evaluation. We propose a regularization-based calibration method that enforces multivariate calibration during training of multivariate distributional regression models using pre-rank functions. We further introduce a novel PCA-based pre-rank that projects predictions onto principal directions of the predictive distribution. Through simulation studies and experiments on 18 real-world multi-output regression datasets, we show that the proposed approach substantially improves multivariate pre-rank calibration without compromising predictive accuracy, and that the PCA pre-rank reveals dependence-structure misspecifications that are not detected by existing pre-ranks.

概率预测多变量校准分布回归预排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。