arXiv:2508.05430cs.CVcs.AI2025-08NeurIPS被引 5

用博弈论方法揭示视觉语言模型相似性背后的复杂跨模态交互。

Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

  • 基于加权Banzhaf指数,分解视觉语言模型的二阶交互影响。
  • 在MS COCO和ImageNet-1k上,二阶解释比一阶方法更准确。
  • 适用于对比CLIP与SigLIP-2等模型的解释能力分析。

语言图像预训练(LIP)使视觉语言模型能够实现零样本分类、定位、多模态检索和语义理解。已有解释方法试图可视化输入图文对在模型相似性输出中的重要性,但主流显著性图仅捕捉一阶贡献,忽略了编码器中固有的复杂跨模态交互。我们提出忠实交互解释方法(FIxLIP),作为统一框架,用于分解视觉语言编码器中的相似性。该方法基于博弈论,采用加权Banzhaf交互指数,在灵活性和计算效率上优于传统的Shapley交互量化框架。从实践角度,我们自然扩展了指向游戏和插入/删除曲线间面积等评价指标,以适应二阶交互解释。在MS COCO和ImageNet-1k基准上的实验表明,如FIxLIP这样的二阶方法优于一阶归因方法。除提供高质量解释外,我们还展示了FIxLIP在比较不同模型(如CLIP与SigLIP-2)方面的实用性。

原文摘要 · Abstract (English)

Language-image pre-training (LIP) enables the development of vision-language models capable of zero-shot classification, localization, multimodal retrieval, and semantic understanding. Various explanation methods have been proposed to visualize the importance of input image-text pairs on the model's similarity outputs. However, popular saliency maps are limited by capturing only first-order attributions, overlooking the complex cross-modal interactions intrinsic to such encoders. We introduce faithful interaction explanations of LIP models (FIxLIP) as a unified approach to decomposing the similarity in vision-language encoders. FIxLIP is rooted in game theory, where we analyze how using the weighted Banzhaf interaction index offers greater flexibility and improves computational efficiency over the Shapley interaction quantification framework. From a practical perspective, we propose how to naturally extend explanation evaluation metrics, such as the pointing game and area between the insertion/deletion curves, to second-order interaction explanations. Experiments on the MS COCO and ImageNet-1k benchmarks validate that second-order methods, such as FIxLIP, outperform first-order attribution methods. Beyond delivering high-quality explanations, we demonstrate the utility of FIxLIP in comparing different models, e.g. CLIP vs. SigLIP-2.

视觉语言模型解释博弈论交互分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。