arXiv:2412.02104cs.CL2024-12综述被引 86

系统梳理多模态大模型可解释性研究,构建三维度分析框架。

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

  • 从数据、模型、训练推理三方面分类现有可解释性方法。
  • 覆盖从词元到嵌入层的多粒度解释,涵盖架构与训练策略。
  • 适合关注AI透明性、可信性的研究人员与开发者参考。

人工智能的快速发展重塑了多个领域,大语言模型(LLMs)与计算机视觉(CV)系统分别推动了自然语言理解与视觉处理的进步。两类技术的融合催生了多模态人工智能,实现跨文本、视觉、音频和视频等模态的丰富理解。多模态大语言模型(MLLMs)成为强大框架,在图像-文本生成、视觉问答与跨模态检索等任务中表现卓越。然而,其复杂性和规模带来显著的可解释性挑战,这对高风险应用场景中的透明度、可信度与可靠性至关重要。本文对MLLMs的可解释性与可解释性提供全面综述,提出一个新型三维度框架:(I) 数据,(II) 模型,(III) 训练与推理。系统分析从词元级到嵌入级的表示可解释性,评估架构分析与设计方法,并探索提升透明度的训练与推理策略。通过比较各类方法,识别其优劣,提出未来研究方向以应对未解挑战。本综述为推进MLLMs的可解释性与透明性提供基础资源,指导研究者与实践者构建更可问责、更鲁棒的多模态AI系统。

原文摘要 · Abstract (English)

The rapid development of Artificial Intelligence (AI) has revolutionized numerous fields, with large language models (LLMs) and computer vision (CV) systems driving advancements in natural language understanding and visual processing, respectively. The convergence of these technologies has catalyzed the rise of multimodal AI, enabling richer, cross-modal understanding that spans text, vision, audio, and video modalities. Multimodal large language models (MLLMs), in particular, have emerged as a powerful framework, demonstrating impressive capabilities in tasks like image-text generation, visual question answering, and cross-modal retrieval. Despite these advancements, the complexity and scale of MLLMs introduce significant challenges in interpretability and explainability, essential for establishing transparency, trustworthiness, and reliability in high-stakes applications. This paper provides a comprehensive survey on the interpretability and explainability of MLLMs, proposing a novel framework that categorizes existing research across three perspectives: (I) Data, (II) Model, (III) Training \& Inference. We systematically analyze interpretability from token-level to embedding-level representations, assess approaches related to both architecture analysis and design, and explore training and inference strategies that enhance transparency. By comparing various methodologies, we identify their strengths and limitations and propose future research directions to address unresolved challenges in multimodal explainability. This survey offers a foundational resource for advancing interpretability and transparency in MLLMs, guiding researchers and practitioners toward developing more accountable and robust multimodal AI systems.

多模态可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。