arXiv:2507.07780cs.CVcs.AI2025-07被引 6

研究图像分类在数据分布偏移下的校准问题,给出实用改进方案。

Where are we with calibration under dataset shift in image classification?

  • 联合使用熵正则化与标签平滑,提升偏移下的原始概率校准效果
  • 少量无关类别异常数据可显著增强校准鲁棒性
  • 微调大模型+集成学习是当前最优校准策略,适合实际应用

我们对真实世界数据分布偏移下图像分类的校准现状进行了广泛研究。通过在八个不同分类任务、多个成像领域上对比多种后处理校准方法及其与训练中校准策略(如标签平滑)的交互作用,发现:(i) 同时采用熵正则化和标签平滑能获得最佳的原始概率校准;(ii) 仅暴露少量语义无关的分布外数据即可使后处理校准器在偏移下最鲁棒;(iii) 针对偏移设计的新型校准方法未必优于简单后处理方法;(iv) 提升偏移下的校准性能常以损害分布内校准为代价。这些结论在随机初始化模型和从基础模型微调的模型上均成立,后者始终表现更优。深入分析集成效应发现:(i) 先校准再集成比反之更有效;(ii) 对于集成模型,分布外数据暴露会恶化分布内-偏移校准权衡;(iii) 集成仍是提升校准鲁棒性的最有效方法之一,结合基础模型微调可达到最佳整体校准效果。

原文摘要 · Abstract (English)

We conduct an extensive study on the state of calibration under real-world dataset shift for image classification. Our work provides important insights on the choice of post-hoc and in-training calibration techniques, and yields practical guidelines for all practitioners interested in robust calibration under shift. We compare various post-hoc calibration methods, and their interactions with common in-training calibration strategies (e.g., label smoothing), across a wide range of natural shifts, on eight different classification tasks across several imaging domains. We find that: (i) simultaneously applying entropy regularisation and label smoothing yield the best calibrated raw probabilities under dataset shift, (ii) post-hoc calibrators exposed to a small amount of semantic out-of-distribution data (unrelated to the task) are most robust under shift, (iii) recent calibration methods specifically aimed at increasing calibration under shifts do not necessarily offer significant improvements over simpler post-hoc calibration methods, (iv) improving calibration under shifts often comes at the cost of worsening in-distribution calibration. Importantly, these findings hold for randomly initialised classifiers, as well as for those finetuned from foundation models, the latter being consistently better calibrated compared to models trained from scratch. Finally, we conduct an in-depth analysis of ensembling effects, finding that (i) applying calibration prior to ensembling (instead of after) is more effective for calibration under shifts, (ii) for ensembles, OOD exposure deteriorates the ID-shifted calibration trade-off, (iii) ensembling remains one of the most effective methods to improve calibration robustness and, combined with finetuning from foundation models, yields best calibration results overall.

模型校准分布偏移集成学习微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。