提出首个评估多模态可解释性的诊断框架,检验模型是否真懂跨模态关系。
GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods

- 构建封闭世界合成数据集,生成数学保证的唯一解释
- 对比真实空间推理与伪快捷方式模型,发现多数方法无法区分
- 适合研究多模态解释性、模型决策机制的学者使用
随着视觉-语言模型的发展,其预测结果需具备可解释性以满足利益相关方需求。然而,可解释性研究未能跟上多模态发展的步伐。现有可解释方法虽能生成跨模态交互解释,但缺乏真实基准来区分真正的跨模态推理(如空间组合)与浅层捷径(如词袋属性匹配)。当前尚不清楚这些方法是否真正捕捉了协同作用,还是仅在模型作为简单特征检测器时虚构推理。本文提出 GridVQA-X,首个专门用于评估跨模态可解释性的诊断框架。不同于自然数据集,GridVQA-X 采用闭世界合成逻辑,生成具有数学保证的独特解释。我们在此控制环境中训练结构相同的两组基线模型:$M_{\text{pure}}$ 学习稳健的空间关系推理,$M_{\text{spur}}$ 被强制依赖跨模态捷径。这种行为差异构成严格测试平台:忠实的解释器必须为两个模型报告不同推理路径。结果表明,广泛使用的解释方法无法区分依赖真实空间推理的模型与利用跨模态捷径的模型,暴露出对真正跨模态协同作用理解的重大缺失,并错误地描绘了多模态模型实际决策过程。
原文摘要 · Abstract (English)
With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily explainable to relevant stakeholders. However, the field of explainability has not kept pace with the multimodal surge. While recent Multimodal Explainable AI (MxAI) methods generate explanations to attribute the interaction between different modalities, current evaluation protocols lack the ground truth required to distinguish between true cross-modal reasoning (e.g., spatial composition) and shallow cross-modal shortcuts (e.g., Bag-of-Words attribute matching). It remains unknown whether MxAI methods faithfully capture synergistic interactions or merely hallucinate reasoning on models acting as simple feature detectors. In this paper, we introduce GridVQA-X, the first diagnostic framework specifically designed to evaluate cross-modal explainability. Unlike natural datasets, GridVQA-X leverages a closed-world synthesis logic to generate unique, mathematically guaranteed explanations. We utilize this controlled environment to train paired ground-truth models on identical architectures: $M_{\text{pure}}$, which learns robust spatial-relational reasoning and $M_{\text{spur}}$, which is structurally forced to rely on cross-modal shortcuts. This behavioral divergence creates a rigorous testbed: a faithful explainer must report distinct reasoning pathways for each model. Our findings reveal that widely used methods fail to distinguish between models relying on genuine spatial-relational reasoning and those exploiting cross-modal shortcuts, highlighting a critical gap in capturing true cross-modal synergy and misrepresenting how multimodal models actually make decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。