arXiv:2508.04427cs.LGcs.AI2025-08综述被引 4

系统梳理多模态模型可解释性研究,指出现有方法的局限与评估标准缺失。

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

  • 分析2020至2024年多模态可解释性论文,聚焦注意力机制应用
  • 发现多数研究集中于视觉-语言模型,且评估方法不统一、缺乏鲁棒性
  • 提出标准化评估与报告建议,推动可解释多模态AI发展

近年来,多模态学习在注意力机制模型的推动下取得显著进展,性能大幅提升。与此同时,可解释人工智能(XAI)需求日益增长,相关研究致力于解析这些模型的决策过程。本系统综述分析了2020年1月至2024年初发表的多模态模型可解释性研究,从模型架构、模态类型、解释算法和评估方法等维度展开。结果显示,多数研究集中于视觉-语言和纯语言模型,注意力机制是最常用的解释手段。然而,现有方法往往难以捕捉模态间复杂交互,且受领域间架构异质性影响。更重要的是,多模态环境下XAI的评估方法普遍缺乏系统性、一致性和对模态特异性认知与上下文因素的考量。为此,我们不仅整合已有研究成果,还结合新兴进展提出综合建议,旨在推动多模态可解释性研究在评估与报告方面的严谨性、透明度与标准化。目标是构建更可解释、可问责、负责任的多模态人工智能系统。

原文摘要 · Abstract (English)

Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks. Parallel to this progress, the demand for explainable artificial intelligence (XAI) has spurred a growing body of research aimed at interpreting the complex decision-making processes of these models. This systematic literature review analyzes research published between January 2020 and early 2024 that focuses on the explainability of multimodal models. Framed within the broader goals of XAI, we examine the literature across multiple dimensions, including model architecture, modalities involved, explanation algorithms and evaluation methodologies. Our analysis reveals that most studies are concentrated on vision-language and language-only models, with attention-based techniques being the most commonly employed for explanation. However, these methods often fall short in capturing the full spectrum of interactions between modalities, a challenge further compounded by the architectural heterogeneity across domains. Importantly, we find that evaluation methods for XAI in multimodal settings are largely non-systematic, lacking consistency, robustness, and consideration for modality-specific cognitive and contextual factors. To address these gaps, we not only synthesize findings from the surveyed works but also incorporate a complementary analysis that integrates recent and emerging advances driving multimodal explainability. Based on these insights, we provide a comprehensive set of recommendations aimed at promoting rigorous, transparent, and standardized evaluation and reporting practices in multimodal XAI research. Our goal is to support future research in more interpretable, accountable, and responsible multimodal AI systems, with explainability at their core.

可解释性多模态注意力机制系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。