arXiv:2511.04160cs.LGstat.ML2025-11被引 1

联合优化深度集成模型的正则化与校准,提升预测性能与不确定性估计。

On Joint Regularization and Calibration in Deep Ensembles

  • 联合调节权重衰减、温度缩放和早停策略
  • 在多个任务上性能匹配或优于独立优化
  • 提出部分重叠验证策略,兼顾评估与数据利用

深度集成是机器学习中提升模型性能和不确定性校准的强大工具。尽管通常通过单独训练和调优各模型形成集成,但已有证据表明联合调优集成整体可带来更好效果。本文研究了联合调节权重衰减、温度缩放和早停对预测性能与不确定性量化的影响。此外,提出一种部分重叠验证策略,作为实现联合评估与最大化数据利用率之间的实用折衷。结果表明,联合调优通常能保持或提升性能,且效果大小在不同任务和指标间存在显著差异。本文强调了个体与联合优化之间的权衡,部分重叠验证策略提供了一种吸引人的实际解决方案。我们相信这些发现为优化深度集成模型的实践者提供了有价值的见解与指导。代码已公开:https://github.com/lauritsf/ensemble-optimality-gap

原文摘要 · Abstract (English)

Deep ensembles are a powerful tool in machine learning, improving both model performance and uncertainty calibration. While ensembles are typically formed by training and tuning models individually, evidence suggests that jointly tuning the ensemble can lead to better performance. This paper investigates the impact of jointly tuning weight decay, temperature scaling, and early stopping on both predictive performance and uncertainty quantification. Additionally, we propose a partially overlapping holdout strategy as a practical compromise between enabling joint evaluation and maximizing the use of data for training. Our results demonstrate that jointly tuning the ensemble generally matches or improves performance, with significant variation in effect size across different tasks and metrics. We highlight the trade-offs between individual and joint optimization in deep ensemble training, with the overlapping holdout strategy offering an attractive practical solution. We believe our findings provide valuable insights and guidance for practitioners looking to optimize deep ensemble models. Code is available at: https://github.com/lauritsf/ensemble-optimality-gap

深度集成正则化校准模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。