arXiv:2509.23665cs.LGcs.AI2025-09被引 2

分析了模型校准方法在不同数据特征下的表现,提升预测可信度。

Calibration Meets Reality: Making Machine Learning Predictions Trustworthy

  • 理论推导校准方法的收敛性与计算复杂度
  • 实验证明在多种数据集上校准性能持续提升
  • 揭示特征质量对校准效果的关键影响,适合做可靠性评估的研究者

后校准方法广泛用于提升机器学习模型概率预测的可靠性。尽管应用广泛,其在不同数据集和模型架构下的理论理解仍不充分,尤其是输入特征质量对校准性能的影响尚未深入研究。本文对Platt缩放和等向回归两种后校准方法进行严格理论分析,推导出收敛保证、计算复杂度边界及有限样本性能指标。通过控制变量的合成实验,探究特征信息量对校准效果的影响。在涵盖多种真实数据集和模型架构的实证评估中,展示了各类场景下校准指标的稳定改进。通过仅使用信息特征与包含噪声维度的完整特征空间对比,揭示不同校准方法的鲁棒性与可靠性。研究结果为根据数据特性和计算约束选择合适校准方法提供了实用指导,弥合了不确定性量化中理论与实践的鸿沟。代码与数据见:https://github.com/Ajwebdevs/calibration-analysis-experiments。

原文摘要 · Abstract (English)

Post-hoc calibration methods are widely used to improve the reliability of probabilistic predictions from machine learning models. Despite their prevalence, a comprehensive theoretical understanding of these methods remains elusive, particularly regarding their performance across different datasets and model architectures. Input features play a crucial role in shaping model predictions and, consequently, their calibration. However, the interplay between feature quality and calibration performance has not been thoroughly investigated. In this work, we present a rigorous theoretical analysis of post-hoc calibration methods, focusing on Platt scaling and isotonic regression. We derive convergence guarantees, computational complexity bounds, and finite-sample performance metrics for these methods. Furthermore, we explore the impact of feature informativeness on calibration performance through controlled synthetic experiments. Our empirical evaluation spans a diverse set of real-world datasets and model architectures, demonstrating consistent improvements in calibration metrics across various scenarios. By examining calibration performance under varying feature conditions utilizing only informative features versus complete feature spaces including noise dimensions, we provide fundamental insights into the robustness and reliability of different calibration approaches. Our findings offer practical guidelines for selecting appropriate calibration methods based on dataset characteristics and computational constraints, bridging the gap between theoretical understanding and practical implementation in uncertainty quantification. Code and experimental data are available at: https://github.com/Ajwebdevs/calibration-analysis-experiments.

模型校准不确定性量化可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。