arXiv:2509.00749cs.CVcs.AI2025-09被引 3

用因果分析揭示视觉自编码器特征的真实触发条件。

Causal Interpretation of Sparse Autoencoder Features in Vision

  • 基于有效感受野,识别真正引发特征激活的图像区域。
  • 发现多个特征需特定上下文组合(如眼鼻共现)才触发。
  • 适合关注模型可解释性与特征语义准确性的研究者。

理解视觉变换器中稀疏自编码器(SAE)特征的语义通常依赖于分析特征激活最高的图像块。然而,自注意力机制会混合全图信息,导致激活块常与特征触发共现而非因果相关。本文提出因果特征解释(CaFE),利用有效感受野(ERF)定位真正驱动特征激活的图像区域。在CLIP-ViT特征中,ERF地图与原始激活地图常不一致,揭示隐藏的上下文依赖关系(如‘咆哮面容’特征需眼鼻同时存在,而不仅是张开的嘴)。插入测试表明,使用CaFE识别的块比按激活强度排序的块更能有效恢复或抑制特征激活。结果表明,CaFE能提供更忠实、语义更精确的SAE特征解释,提示仅依赖激活位置可能导致误读。

原文摘要 · Abstract (English)

Understanding what sparse auto-encoder (SAE) features in vision transformers truly represent is usually done by inspecting the patches where a feature's activation is highest. However, self-attention mixes information across the entire image, so an activated patch often co-occurs with-but does not cause-the feature's firing. We propose Causal Feature Explanation (CaFE), which leverages Effective Receptive Field (ERF). We consider each activation of an SAE feature to be a target and apply input-attribution methods to identify the image patches that causally drive that activation. Across CLIP-ViT features, ERF maps frequently diverge from naive activation maps, revealing hidden context dependencies (e.g., a "roaring face" feature that requires the co-occurrence of eyes and nose, rather than merely an open mouth). Patch insertion tests confirm that CaFE more effectively recovers or suppresses feature activations than activation-ranked patches. Our results show that CaFE yields more faithful and semantically precise explanations of vision-SAE features, highlighting the risk of misinterpretation when relying solely on activation location.

可解释性视觉模型因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。