arXiv:2502.17516cs.LGcs.AI2025-02综述被引 41

系统梳理多模态大模型可解释性方法,填补与语言模型解释力的差距。

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

  • 将语言模型解释技术适配到多模态模型中
  • 提出多模态可解释性方法的结构化分类体系
  • 揭示跨模态系统与单模态模型的本质差异

基础模型的兴起改变了机器学习研究格局,推动了对其内部机制的理解和更高效、可靠的可控应用。尽管在大语言模型(LLMs)的可解释性方面已取得显著进展,但多模态基础模型(MMFMs),如对比视觉-语言模型、生成式视觉-语言模型和文生图模型,相较于单模态框架提出了独特的可解释性挑战。尽管已有初步研究,但在可解释性水平上,LLMs与MMFMs之间仍存在显著差距。本综述聚焦两个核心方面:(1) LLM可解释性方法向多模态模型的迁移适配;(2) 理解单模态语言模型与跨模态系统之间的机制差异。通过系统回顾现有MMFM分析技术,我们提出了一个结构化的可解释性方法分类体系,比较了单模态与多模态架构中的洞察,并指出了关键研究空白。

原文摘要 · Abstract (English)

The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for better control. While significant progress has been made in interpreting Large Language Models (LLMs), multimodal foundation models (MMFMs) - such as contrastive vision-language models, generative vision-language models, and text-to-image models - pose unique interpretability challenges beyond unimodal frameworks. Despite initial studies, a substantial gap remains between the interpretability of LLMs and MMFMs. This survey explores two key aspects: (1) the adaptation of LLM interpretability methods to multimodal models and (2) understanding the mechanistic differences between unimodal language models and crossmodal systems. By systematically reviewing current MMFM analysis techniques, we propose a structured taxonomy of interpretability methods, compare insights across unimodal and multimodal architectures, and highlight critical research gaps.

可解释性多模态基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。