发现CLIP模型中特征的感知范围差异,揭示其对预测影响机制。
Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
- 提出上下文依赖度(CDS)量化特征在图像中的局部或全局响应范围。
- 不同范围特征对CLIP预测和置信度有系统性影响,局部特征更稳定。
- 为理解CLIP内部表示提供新视角,适合关注模型可解释性的研究者。
稀疏自编码器(SAEs)已成为解析CLIP视觉编码器内部表征的强大工具,但现有分析主要聚焦于单个特征的语义意义。本文引入信息范围作为互补的可解释性维度,描述特征聚合视觉证据的广度,从局部的、块级线索到全局的、图像级信号。我们观察到,某些SAE特征在空间扰动下保持一致响应,而另一些则随微小输入变化而显著漂移,表明其基础范围存在根本差异。为此,我们提出上下文依赖度(CDS)来量化该特性,区分位置稳定的局部范围特征与位置变化的全局范围特征。实验表明,不同信息范围的特征对CLIP预测及其置信度产生系统性影响。这些发现确立了信息范围作为理解CLIP表征的关键新维度,并提供了对SAE衍生特征更深层的诊断视角。
原文摘要 · Abstract (English)
Sparse Autoencoders (SAEs) have emerged as a powerful tool for interpreting the internal representations of CLIP vision encoders, yet existing analyses largely focus on the semantic meaning of individual features. We introduce information scope as a complementary dimension of interpretability that characterizes how broadly an SAE feature aggregates visual evidence, ranging from localized, patch-specific cues to global, image-level signals. We observe that some SAE features respond consistently across spatial perturbations, while others shift unpredictably with minor input changes, indicating a fundamental distinction in their underlying scope. To quantify this, we propose the Contextual Dependency Score (CDS), which separates positionally stable local scope features from positionally variant global scope features. Our experiments show that features of different information scopes exert systematically different influences on CLIP's predictions and confidence. These findings establish information scope as a critical new axis for understanding CLIP representations and provide a deeper diagnostic view of SAE-derived features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。