arXiv:2503.12605cs.CV2025-03综述被引 203

系统梳理多模态思维链推理,助力模型像人一样一步步理解图像视频等信息。

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

  • 构建多模态思维链的统一框架,融合图像、视频、语音等多源信息逐步推理。
  • 归纳10+类方法与应用场景,涵盖医疗、自动驾驶、机器人等关键领域。
  • 为迈向通用人工智能提供方向,适合关注多模态大模型的研究者参考。

将人类式分步思考(思维链)的优势拓展至多模态场景,多模态思维链(MCoT)推理近年受到广泛关注,尤其在多模态大语言模型(MLLMs)中的融合应用。现有研究设计多种方法与创新推理范式,应对图像、视频、语音、音频、3D及结构化数据等不同模态的独特挑战,在机器人、医疗、自动驾驶和多模态生成等领域取得广泛成功。然而,MCoT仍面临显著挑战与机遇,亟需深入研究,而当前该领域尚无系统性综述。为此,本文首次提出MCoT推理的系统性综述,阐明相关基础概念与定义,构建全面分类体系,并从多角度深入分析现有方法在不同应用场景的表现。同时,揭示现存挑战并展望未来研究方向,旨在推动多模态通用人工智能的发展。

原文摘要 · Abstract (English)

By extending the advantage of chain-of-thought (CoT) reasoning in human-like step-by-step processes to multimodal contexts, multimodal CoT (MCoT) reasoning has recently garnered significant research attention, especially in the integration with multimodal large language models (MLLMs). Existing MCoT studies design various methodologies and innovative reasoning paradigms to address the unique challenges of image, video, speech, audio, 3D, and structured data across different modalities, achieving extensive success in applications such as robotics, healthcare, autonomous driving, and multimodal generation. However, MCoT still presents distinct challenges and opportunities that require further focus to ensure consistent thriving in this field, where, unfortunately, an up-to-date review of this domain is lacking. To bridge this gap, we present the first systematic survey of MCoT reasoning, elucidating the relevant foundational concepts and definitions. We offer a comprehensive taxonomy and an in-depth analysis of current methodologies from diverse perspectives across various application scenarios. Furthermore, we provide insights into existing challenges and future research directions, aiming to foster innovation toward multimodal AGI.

多模态思维链大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。