以iERF为核心,统一视觉模型局部、全局与机制解释
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models

- 用点式特征向量和实例化感受野统一分析单元
- 在多种模型上优于基线,可解释分散的SAE特征
- 适合研究模型决策路径与对抗样本的可解释性
现代视觉模型虽准确率高,但其证据来源、编码内容及内部计算如何整合仍缺乏统一解释。本文提出以实例化有效感受野(iERF)为中心的统一框架,将局部、全局与机制解释统一到点式特征向量(PFV)之上。局部层面,通过共享比率分解(SRD)将每个PFV表示为上游特征的混合,并传播iERF生成类判别显著图,实现高分辨率、激活忠实且对扰动鲁棒的解释。全局层面,提出概念锚定特征解释(CAFE),以iERF作为语义标签,将抽象潜在向量与可验证的像素级证据关联,解决Transformer中早期自注意力导致的非局部稀疏自动编码器特征问题。针对深度表征如何组合,提出层间概念图与层间概念归因(ICAT),量化概念间影响并隔离层对;插入/删除协议验证积分梯度是最忠实的实现方式。实证表明,在ResNet50、VGG16与ViTs上,该框架在保真度与鲁棒性上均优于基线,成功解析分散的SAE特征,并揭示正确分类、误分类及对抗样本中的主导概念路径。基于iERF,本方法提供了从像素到概念再到决策的连贯、有证据支持的映射。
原文摘要 · Abstract (English)
Modern vision models achieve remarkable accuracy, but explaining where evidence arises, what the model encodes, and how internal computations assemble that evidence remains fragmented. We introduce an iERF-centric framework that unifies local, global, and mechanistic interpretability around a single analysis unit: the pointwise feature vector (PFV) paired with its instance-specific Effective Receptive Field (iERF). On the local side, Sharing Ratio Decomposition (SRD) expresses each PFV as a mixture of upstream PFVs via sharing ratios and propagates iERFs to construct class-discriminative saliency maps. SRD yields high-resolution, activation-faithful explanations, is robust to targeted manipulation and noise, and remains activation-agnostic across common nonlinearities. For the global view, we introduce Concept-Anchored Feature Explanation (CAFE), which utilizes the iERF as a semantic label, grounding abstract latent vectors in verifiable pixel-level evidence. With CAFE, we address the challenge of non-localized sparse autoencoder latents--especially in Transformers, where early self-attention mixes distant context. To answer how representations are composed through depth, we propose the Interlayer Concept Graph with Interlayer Concept Attribution (ICAT), which quantifies concept-to-concept influence while isolating layer pairs; an interlayer insertion, deletion protocol identifies Integrated Gradients as the most faithful instantiation. Empirically, across ResNet50, VGG16, and ViTs, our framework outperforms baselines in both fidelity and robustness, successfully interprets dispersed SAE features, and exposes dominant concept routes in correct, misclassified, and adversarial cases. Grounded in iERFs, our approach provides a coherent, evidence-backed map from pixels to concepts to decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。