提出新评估方法,精准衡量视觉语言模型的跨模态协同能力。
Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability

- 基于谢泼德交互指数设计新度量,严格分离模态间协同效应。
- 实测显示现有解释方法严重高估视觉显著性,低估跨模态协同。
- 适合关注模型可解释性与高风险应用安全审计的研究者。
视觉-语言模型(VLM)将复杂视觉输入映射到语义空间,但当前对跨模态推理的解释依赖于单模态扰动指标。我们发现该范式存在局限:多模态数据集包含语言先验和模态偏见,导致VLM常表现出跨模态冗余,仅凭文本即可回答视觉问题。因此,单模态指标会错误惩罚忠实的解释器,引发评估崩溃(肯德尔τ = -0.06)。为此,我们提出协同忠实度($ℓ_{syn}$),一种基于谢泼德交互指数的可扩展度量,能严格隔离模态间的联合哈桑尼红利,作为高精度代理指标(ρ = 0.92),并实现24倍计算加速。在3个基准数据集、3种VLM架构上评估8种XAI方法,结果表明现有解释方法过度依赖视觉显著性,显著弱于适配的注意力机制方法在捕捉真实跨模态协同方面的能力。通过解耦视觉合理性与跨模态忠实度,本工作为高风险部署中安全审计VLM推理提供了严谨评估框架。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) map complex visual inputs to semantic spaces, but interpreting the cross-modal reasoning of VLMs currently relies on post-hoc explainers evaluated via unimodal perturbation metrics. We expose a limitation in this paradigm: because multimodal datasets contain language priors and modality biases, VLMs frequently exhibit cross-modal redundancy, allowing them to answer visual queries using text alone. Consequently, unimodal metrics penalize faithful explainers, triggering an evaluation collapse where visual and textual rankings fundamentally contradict each other. %(Kendall's $τ= -0.06$). To resolve this, we introduce Synergistic Faithfulness ($\mathcal{F}_{syn}$), a scalable metric rooted in the Shapley Interaction Index that strictly isolates the joint Harsanyi dividend between modalities, serving as a highly accurate surrogate ($ρ= 0.92$) while achieving a $24\times$ computational speedup. Evaluating 8 distinct XAI methods across 3 VLM architectures and 3 benchmark datasets, reveals that explainers proposed for VLMs heavily over-index on visual salience and significantly underperform adapted attention-based methods in capturing true cross-modal synergy. By decoupling visual plausibility from cross-modal faithfulness, this work provides a rigorous evaluation framework required to safely audit VLM reasoning in high-stakes deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。