提出首个量化多模态大模型决策依赖的归因方法,解决'哪个模态决定结论'的关键问题。
Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

- 基于反事实生成与沙普利值,量化图像和文本对预测的贡献
- 在98%的模拟案例中准确识别决策主导模态,优于现有方法
- 适合安全关键场景下多模态模型的可信度审计
多模态大语言模型(MLLMs)通过融合图像与文本信息支持高风险决策。现有可解释性方法仅能定位重要图像区域或文本词元,无法回答核心问题:哪个模态驱动了预测?这可能导致模型虽正确输出却依赖错误证据,掩盖捷径学习与不安全推理。本文将模态归因定义为多模态基础模型的补充可解释性目标,提出首个量化模态贡献的框架——反事实模态归因(CMA)。CMA利用耦合扩散先验生成仅图像、仅文本及联合多模态反事实样本,并通过基于沙普利值的合作博弈理论转化为严谨的模态归因分数。在可控合成基准(已知模态依赖真值)和真实世界多模态临床数据集上评估,CMA在98%的控制案例中正确识别决策主导模态,持续超越基线方法,揭示出仅凭预测准确率无法发现的跨模态推理失败。结果表明,模态归因是超越特征归因的互补可解释维度,为安全关键应用中的多模态基础模型审计提供了原则性框架。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental question: which modality drives a prediction? Consequently, a model may produce the correct output while relying on the wrong source of evidence, masking shortcut learning and unsafe reasoning. We formulate modality attribution as a complementary explainability objective for multimodal foundation models and propose Counterfactual Modality Attribution (CMA), the first framework for quantifying modality-level contributions in MLLMs. CMA generates image-only, text-only, and joint multimodal counterfactuals using coupled diffusion priors and converts them into principled modality attribution scores through a cooperative game-theoretic formulation based on Shapley values. We evaluate CMA on controlled synthetic benchmarks with known ground-truth modality reliance and on a real-world multimodal clinical dataset. CMA correctly identifies the decision-driving modality in 98% of controlled cases and consistently outperforms baselines, revealing failures of cross-modal reasoning that remain invisible to predictive accuracy alone. Our results establish modality attribution as a complementary dimension of explainability beyond feature attribution, providing a principled framework for auditing multimodal foundation models in safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。