arXiv:2412.14056cs.CVcs.AI2024-12综述被引 26

梳理多模态可解释AI发展历程,助你理解AI决策为何可信。

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

  • 按四个时代分类多模态可解释AI方法,从传统模型到大语言模型。
  • 总结常用评估指标与数据集,揭示当前研究瓶颈。
  • 适合关注AI透明性、公平性的研究者与工程师参考。

人工智能在算力提升和海量数据推动下快速发展,但其“黑箱”特性也带来了可解释性挑战。为增强人类对AI决策的理解与信任,可解释AI(XAI)应运而生。在多模态数据融合与复杂推理场景中,多模态可解释AI(MXAI)整合多种模态完成预测与解释任务。随着大语言模型(LLMs)在自然语言处理中的突破,其复杂性进一步加剧了MXAI的难题。本文从历史视角回顾MXAI的发展,将其分为四个阶段:传统机器学习、深度学习、判别式基础模型和生成式大语言模型。同时综述了MXAI研究中使用的评估指标与数据集,并讨论未来挑战与方向。相关项目已开源:https://github.com/ShilinSun/mxai_review。

原文摘要 · Abstract (English)

Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challenges in interpreting the "black-box" nature of AI models. To address these concerns, eXplainable AI (XAI) has emerged with a focus on transparency and interpretability to enhance human understanding and trust in AI decision-making processes. In the context of multimodal data fusion and complex reasoning scenarios, the proposal of Multimodal eXplainable AI (MXAI) integrates multiple modalities for prediction and explanation tasks. Meanwhile, the advent of Large Language Models (LLMs) has led to remarkable breakthroughs in natural language processing, yet their complexity has further exacerbated the issue of MXAI. To gain key insights into the development of MXAI methods and provide crucial guidance for building more transparent, fair, and trustworthy AI systems, we review the MXAI methods from a historical perspective and categorize them across four eras: traditional machine learning, deep learning, discriminative foundation models, and generative LLMs. We also review evaluation metrics and datasets used in MXAI research, concluding with a discussion of future challenges and directions. A project related to this review has been created at https://github.com/ShilinSun/mxai_review.

可解释AI多模态大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。