研究视觉定位模型在语义不匹配时的错误原因,发现嵌入方向性不是主因。
Investigating Anisotropy in Visual Grounding under Controlled Counterfactual Perturbations

- 通过控制相似度的反事实提示生成,精细分析模型对语义变化的响应。
- 实验表明嵌入相似度与定位近似行为无显著关联,跨模型结果一致。
- 强调需深入探究嵌入空间的几何特性,提升模型可解释性与可靠性。
视觉定位基准通常假设指代表达描述的对象始终存在于图像中,因此模型很少在语义不匹配的场景下被评估。此类情况下,模型常表现出近似行为,仅满足表达的部分内容(如保留原对象而忽略修改的上下文线索),导致结果不可靠并引发可解释性担忧。本文从机制可解释性视角出发,探究嵌入方向性是否导致反事实失败。提出一种受控相似度的反事实标题生成方法,在预定义的嵌入相似度区间内系统扰动对象或上下文成分,实现对定位行为随对齐程度变化的细粒度分析。在两种嵌入几何差异显著的Transformer模型(基于BERT的TransVG和基于CLIP的SwimVG)上进行实验,结果显示余弦相似度与近似行为之间无明显相关性。该结果表明,仅靠嵌入方向性无法解释反事实错误,提升鲁棒性需进一步研究嵌入空间的更细粒度几何属性。
原文摘要 · Abstract (English)
Visual Grounding benchmarks assume that the object described by a referring expression is always present in the image, and grounding models are therefore rarely evaluated under semantically mismatched captions. In such cases, models frequently exhibit approximation behavior, producing a plausible bounding box that satisfies only part of the expression (\eg, preserving the original object while ignoring modified contextual cues). Because mismatched captions represent realistic edge cases, this behavior compromises reliability and raises concerns from an explainability perspective. Identifying its underlying causes is thus essential for improving model faithfulness and interpretability. Adopting a mechanistic interpretability viewpoint, this work examines whether embedding anisotropy contributes to counterfactual failures. A similarity-controlled counterfactual caption generation protocol is introduced to systematically perturb object or contextual components within predefined embedding similarity intervals, enabling a fine-grained analysis of grounding behavior as a function of alignment. Experiments on two Transformer-based models with markedly different embedding geometries (BERT-based TransVG and CLIP-based SwimVG) reveal no meaningful correlation between cosine similarity and approximation. These findings suggest that anisotropy alone does not account for counterfactual errors, and that robustness requires investigating finer-grained geometric properties of the embedding space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。