视觉模型中隐含的隐喻现象破坏了注意力解释性,导致部件分析失效。
Metonymy in vision models undermines attention-based interpretability

- 通过部件属性标注检验模型局部性假设,发现现代ViT存在强跨部件信息泄露。
- 实验显示各部件编码整个物体信息,形成视觉隐喻,使注意力解释不可靠。
- 两阶段设计可消除泄露,提升属性驱动的部件发现性能,适合可解释性研究者。
基于部件的推理是让计算机视觉模型直接关注与下游任务相关的物体部件的经典策略。在深度学习背景下,这也能通过在标准黑箱模型的潜在表示上添加以部件为中心的注意力机制来实现可解释性设计。该方法依赖于局部性假设:物体部件的潜在表示主要编码对应图像区域的信息。本文通过部件属性标注全面评估了视觉模型中的部件内信息泄露,结果表明现代预训练视觉变压器违背了局部性假设,表现出强烈的部件内泄露——每个部件编码整个物体的信息,形成一种视觉隐喻,损害了基于注意力的可解释性方法对部件推理的忠实性,最终使其失去可解释性。此外,我们提出一种两阶段方法作为上限,通过设计防止泄露;并证明这种内在解耦的特征提取能显著提升多种任务下的属性驱动部件发现效果,证实了部件内泄露的实际影响。我们的研究揭示了影响部件表示可解释性的被忽视问题,尤其针对依赖部件中心概念的类别感知模型(CBMs),指出两阶段方法是缓解该问题的有前景方向。
原文摘要 · Abstract (English)
Part-based reasoning is a classical strategy to make a computer vision model directly focus on the object parts that are relevant to the downstream task. In the context of deep learning, this also serves to improve by-design interpretability, often by using part-centric attention mechanisms on top of a latent image representation provided by a standard, black-box model. This approach is based on a locality assumption: that the latent representation of an object part encodes primarily information about the corresponding image region. In this work, we test this basic assumption, measuring intra-object leakage in vision models using part-based attribute annotations. Through a comprehensive experimental evaluation, we show that modern pretrained vision transformers violate the locality assumption and exhibit a strong intra-object leakage, in which each part encodes information from the whole object, a visual metonymy that compromises the faithfulness of attention-based interpretable-by-design methods for part-based reasoning, ultimately rendering them uninterpretable. In addition, we establish an upper bound using a two-stage approach that prevents leakage by design. We then show that this inherently disentangled feature extraction improves attribute-driven part discovery on a variety of tasks, confirming the practical impact of intra-object leakage. Our results uncover a neglected issue affecting the interpretability of part-based representations, such as those in CBMs relying on part-centric concepts, highlighting that two-stage approaches offer a promising way to mitigate it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。