arXiv:2601.10836cs.CV2026-01

训练方式影响模型的异常检测能力,好成绩未必带来更好泛化判断。

One Model, Many Behaviors: Training-Induced Effects on Out-of-Distribution Detection

  • 用相同模型测试21种检测方法,发现训练策略决定检测效果
  • 当准确率超过一定水平后,异常检测能力反而下降
  • 没有万能检测器,需根据训练方式选合适方法

开放世界中,异常样本检测对机器学习系统的鲁棒性至关重要。尽管异常检测技术持续进步,但其与现代训练流程(以提升分布内准确率和泛化能力为目标)之间的关系仍不明确。我们通过全面的实证研究探索这一关联。固定架构为广泛使用的ResNet-50,基于56个通过不同训练策略在ImageNet上训练的模型,评估21种前沿后处理型异常检测方法在8个异常数据集上的表现。结果表明,高分布内准确率并不必然带来更好的异常检测性能:异常检测能力随准确率提升先上升后下降,呈现非单调关系。此外,训练策略、检测方法选择与最终性能高度相关,说明不存在普适最优的检测方案。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection is crucial for deploying robust and reliable machine-learning systems in open-world settings. Despite steady advances in OOD detectors, their interplay with modern training pipelines that maximize in-distribution (ID) accuracy and generalization remains under-explored. We investigate this link through a comprehensive empirical study. Fixing the architecture to the widely adopted ResNet-50, we benchmark 21 post-hoc, state-of-the-art OOD detection methods across 56 ImageNet-trained models obtained via diverse training strategies and evaluate them on eight OOD test sets. Contrary to the common assumption that higher ID accuracy implies better OOD detection performance, we uncover a non-monotonic relationship: OOD performance initially improves with accuracy but declines once advanced training recipes push accuracy beyond the baseline. Moreover, we observe a strong interdependence between training strategy, detector choice, and resulting OOD performance, indicating that no single method is universally optimal.

异常检测模型训练深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。