提出新方法打破特征空间语义霸权,让模型更敏感地识别结构异常的离群样本。
Breaking Semantic Hegemony: Decoupling Principal and Residual Subspaces for Generalized OOD Detection
- 通过正交分解分离语义与结构残差,解耦特征空间
- 在CIFAR和ImageNet上将FPR95从31.3%降至2.3%
- 适合需要高鲁棒性检测的工业级视觉系统
尽管基于特征的后处理方法在离群检测(OOD)方面取得显著进展,我们发现现有最先进模型存在反直觉的‘简约悖论’:对语义细微差异的离群样本敏感,却对结构迥异但语义简单的样本或高频传感器噪声表现出严重几何盲区。我们归因于深度特征空间中的语义霸权,并通过神经坍缩理论揭示其数学本质——主子空间的高方差导致谱集中偏差,数值上掩盖了残差子空间中本应显著的结构分布偏移信号。为此,我们提出D-KNN,一种无需训练、即插即用的几何解耦框架。该方法利用正交分解显式分离语义成分与结构残差,并引入双空间校准机制,重激活模型对微弱残差信号的敏感性。大量实验表明,D-KNN有效打破语义霸权,在CIFAR与ImageNet基准上建立新SOTA。尤其在解决简约悖论时,将FPR95从31.3%降至2.3%;面对高斯噪声等传感器故障时,检测性能(AUROC)从79.7%提升至94.9%。
原文摘要 · Abstract (English)
While feature-based post-hoc methods have made significant strides in Out-of-Distribution (OOD) detection, we uncover a counter-intuitive Simplicity Paradox in existing state-of-the-art (SOTA) models: these models exhibit keen sensitivity in distinguishing semantically subtle OOD samples but suffer from severe Geometric Blindness when confronting structurally distinct yet semantically simple samples or high-frequency sensor noise. We attribute this phenomenon to Semantic Hegemony within the deep feature space and reveal its mathematical essence through the lens of Neural Collapse. Theoretical analysis demonstrates that the spectral concentration bias, induced by the high variance of the principal subspace, numerically masks the structural distribution shift signals that should be significant in the residual subspace. To address this issue, we propose D-KNN, a training-free, plug-and-play geometric decoupling framework. This method utilizes orthogonal decomposition to explicitly separate semantic components from structural residuals and introduces a dual-space calibration mechanism to reactivate the model's sensitivity to weak residual signals. Extensive experiments demonstrate that D-KNN effectively breaks Semantic Hegemony, establishing new SOTA performance on both CIFAR and ImageNet benchmarks. Notably, in resolving the Simplicity Paradox, it reduces the FPR95 from 31.3% to 2.3%; when addressing sensor failures such as Gaussian noise, it boosts the detection performance (AUROC) from a baseline of 79.7% to 94.9%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。