arXiv:2506.13060cs.AIcs.LG2025-06被引 8

提出多模态解释新范式,解决单一模态解释的误导问题。

Rethinking Explainability in the Era of Multimodal AI

  • 基于模态影响、协同忠实与统一稳定性三原则设计多模态解释方法。
  • 实验证明单模态解释会遗漏跨模态依赖,导致决策误判。
  • 适合高风险场景下的模型可信性评估与安全部署研究者。

多模态人工智能系统(联合训练于文本、时间序列、图结构和图像等异构数据)已在高风险应用中广泛应用并取得卓越性能,但透明且准确的解释算法对确保其安全部署与用户信任至关重要。然而,现有解释技术大多局限于单模态,仅生成各模态独立的特征归因、概念或电路轨迹,无法捕捉跨模态交互。本文指出,此类单模态解释系统性地误判并遗漏驱动多模态模型决策的跨模态影响,呼吁停止依赖它们来解释多模态模型。为此,我们提出了基于模态的三大核心原则:格兰杰式模态影响(通过受控消融量化移除某一模态对另一模态解释的影响)、协同忠实性(解释需反映模态组合时模型的预测能力)与统一稳定性(解释在小规模跨模态扰动下保持一致)。这一向多模态解释的转变将有助于揭示隐藏捷径、缓解模态偏见、提升模型可靠性,并增强高风险场景中的安全性。

原文摘要 · Abstract (English)

While multimodal AI systems (models jointly trained on heterogeneous data types such as text, time series, graphs, and images) have become ubiquitous and achieved remarkable performance across high-stakes applications, transparent and accurate explanation algorithms are crucial for their safe deployment and ensure user trust. However, most existing explainability techniques remain unimodal, generating modality-specific feature attributions, concepts, or circuit traces in isolation and thus failing to capture cross-modal interactions. This paper argues that such unimodal explanations systematically misrepresent and fail to capture the cross-modal influence that drives multimodal model decisions, and the community should stop relying on them for interpreting multimodal models. To support our position, we outline key principles for multimodal explanations grounded in modality: Granger-style modality influence (controlled ablations to quantify how removing one modality changes the explanation for another), Synergistic faithfulness (explanations capture the model's predictive power when modalities are combined), and Unified stability (explanations remain consistent under small, cross-modal perturbations). This targeted shift to multimodal explanations will help the community uncover hidden shortcuts, mitigate modality bias, improve model reliability, and enhance safety in high-stakes settings where incomplete explanations can have serious consequences.

多模态可解释性模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。