arXiv:2606.03795cs.CV2026-06被引 1

分析视觉模型如何改变图像频谱信息,发现中间层影响最大。

Beyond Compression: Quantifying Spectral Accessibility in Vision Representations

论文配图:Beyond Compression: Quantifying Spectral Accessibility in Vision Representations
图 1 · 摘自论文原文
  • 用频谱可恢复性衡量视觉表示中的频率信息变化
  • 中间层频谱可访问性最高,最终输出因架构而异
  • 适合研究模型表征机制或对比不同视觉架构的读者

视觉语言模型通过学习的投影层将视觉特征映射到共享嵌入空间,但这些变换如何改变视觉信息结构尚不明确。本研究通过傅里叶能量的线性可恢复性,衡量视觉表示在空间频率上的可访问性。为排除维度压缩的干扰,提出残差频谱损失(RSL),与维度匹配的随机投影基线对比。实验使用参数冻结的预训练模型,避免优化偏差。结果表明,CLIP和DINOv2在ImageNet与MS-COCO数据集上均表现出一致的频率依赖性变化:频谱可访问性随网络深度呈非单调变化,中间层达到峰值后下降。最终表示差异显著:CLIP的投影为频谱中性,变化主要由压缩导致;而DINOv2的[CLS]池化引入了结构化的频谱损失。研究揭示中间层和池化机制是现代视觉编码器频谱变换的主要驱动因素。

原文摘要 · Abstract (English)

Vision-language models map visual features into a shared embedding space through learned projection layers, yet it remains unclear how these transformations alter the structure of visual information. This study examines changes in representation through spatial-frequency accessibility, measured by the linear recoverability of band-limited Fourier energy from model representations. To isolate effects beyond dimensionality reduction, we introduce Residual Spectral Loss (RSL), which evaluates changes relative to a dimension-matched random projection baseline. To reduce confounding effects from optimization, the analysis uses pretrained models with all parameters frozen. The experimental results show consistent frequency-dependent changes in accessibility across CLIP and DINOv2 on ImageNet and MS-COCO datasets. Spectral accessibility follows a non-monotonic trajectory across depth, peaking at intermediate layers before decreasing toward the output representation. The final transformation differs across architectures: CLIP's learned projection is spectrally neutral, with changes explained by compression, whereas DINOv2's [CLS] pooling induces a structured loss across the spectrum. These findings identify intermediate layers and pooling mechanisms as primary drivers of spectral transformation in modern vision encoders.

视觉表征频谱分析模型机制CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。