追踪视觉模型中间层的表示演变,提升异常检测与分类性能。
Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification

- 分析模型各层间表示的动态演化路径,分离共性迁移与输入特异性变化。
- 在152组对比中,93%情况下降低错误率95%,尤其对视觉扰动和语义偏离有效。
- 适用于各类模型架构,为模型可靠性提供通用信号,适合关注鲁棒性的研究者。
视觉模型的表示并非一蹴而就,每一层都在不断修正前一层结果。我们探究这种计算路径是否包含最终表示所丢弃的信息,并验证其对分布外(OOD)检测与图像分类在干净数据及分布偏移数据上的帮助。不同于将中间层视为独立快照的方法,我们保留样本身份,研究连续层间的变换关系。通过分离类别一致的迁移与输入特异性创新,以及坐标移动与关系重组,发现这些路径在监督、自监督、视觉-语言、层次化和卷积编码器中均表现出强样本特异性连续性和架构特定的深度模式,且跨数据集重复出现。实际应用中,仅使用内部数据(ID)的过渡惊喜度得分可有效补充现有检测器,在131/152个非饱和对比中降低FPR95;在视觉干扰和语义远离的偏移下增益最大,多数探测器在近似OOD场景也保持正向提升。冻结更新探针在71/72个干净模型-数据集组合中表现更优,而偏移数据收益随架构与扰动类型变化。因此,计算路径提供了一种广泛适用的可靠性信号,其价值由模型结构与遭遇的分布偏移共同决定。
原文摘要 · Abstract (English)
Vision models do not form a representation at once; each block revises it. We ask whether the resulting computation path contains evidence that the final representation discards, and whether that evidence improves OOD detection and image classification on clean and shifted data. Unlike approaches that treat intermediate layers as separate snapshots, we retain sample identity across depth and study the transformations connecting successive states. We separate class-coherent transport from input-specific innovation, and coordinate movement from relational reorganization. Across supervised, self-supervised, vision--language, hierarchical, and convolutional encoders, these paths show strong sample-specific continuity and architecture-specific depth profiles that recur across datasets. They are also practically useful. An ID-only transition-surprise score complements strong final-state detectors, reducing FPR95 in 131/152 non-saturated comparisons on a balanced OpenOOD grid; gains are largest for visually disruptive and semantically far shifts, and remain positive on near-OOD for most detectors. Frozen update probes improve 71/72 clean model--dataset cases, while shifted-data gains vary with architecture and corruption type. Computation paths therefore provide a broadly useful reliability signal whose value is determined jointly by model organization and the shift encountered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。