arXiv:2510.05782cs.CV2025-10NeurIPS被引 6

发现预训练模型中间层能显著提升分布外检测效果

Mysteries of the Deep: Role of Intermediate Representations in Out of Distribution Detection

  • 利用熵准则自动筛选对分布外检测最有效的中间层
  • 相比现有方法,远分布外检测准确率提升10%,近分布外超7%
  • 无需标注数据,适用于多种模型架构与训练目标

分布外(OOD)检测对机器学习模型在真实场景中的可靠部署至关重要。然而,现有方法通常将大型预训练模型视为黑箱编码器,仅依赖其最终层表征进行检测。本文挑战这一常规做法,揭示预训练模型的中间层——受残差连接影响,逐步变换输入投影——可编码出丰富且多样的分布偏移信号。为利用各层间的表征多样性,我们提出一种基于熵的准则,在无需访问分布外数据的训练无关设定下,自动识别提供互补信息的层。实验表明,选择性融合这些中间表征,可在多种模型架构与训练目标下,使远分布外检测准确率提升最高达10%,近分布外检测提升超过7%。研究揭示了新的OOD检测方向,并阐明不同训练目标与模型结构对置信度基方法的影响。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection is essential for reliably deploying machine learning models in the wild. Yet, most methods treat large pre-trained models as monolithic encoders and rely solely on their final-layer representations for detection. We challenge this wisdom. We reveal the \textit{intermediate layers} of pre-trained models, shaped by residual connections that subtly transform input projections, \textit{can} encode \textit{surprisingly rich and diverse signals} for detecting distributional shifts. Importantly, to exploit latent representation diversity across layers, we introduce an entropy-based criterion to \textit{automatically} identify layers offering the most complementary information in a training-free setting -- \textit{without access to OOD data}. We show that selectively incorporating these intermediate representations can increase the accuracy of OOD detection by up to \textbf{$10\%$} in far-OOD and over \textbf{$7\%$} in near-OOD benchmarks compared to state-of-the-art training-free methods across various model architectures and training objectives. Our findings reveal a new avenue for OOD detection research and uncover the impact of various training objectives and model architectures on confidence-based OOD detection methods.

分布外检测中间表征无训练方法模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。