系统梳理多模态推理新范式,从语言主导到动态交互
Mind with Eyes: from Language Reasoning to Multimodal Reasoning
- 按视觉辅助程度分为单向感知与协同推理两类
- 提出动作生成与状态更新机制,实现模态间动态交互
- 适合关注多模态智能体与跨模态推理的研究者
语言模型近期已迈向推理领域,但唯有通过多模态推理才能真正实现更全面、类人的认知能力。本综述系统梳理了近年多模态推理方法,将其分为两类:以语言为中心的多模态推理(包括单次视觉感知与主动视觉感知,视觉主要辅助语言推理)和协作式多模态推理(在推理过程中生成动作并更新状态,实现模态间动态交互)。进一步分析了技术演进路径,讨论内在挑战,并介绍关键基准任务与评估指标。最后从视觉-语言推理向全模态推理、多模态推理向多模态智能体两个方向展望未来研究。本综述旨在提供结构化视角,推动多模态推理研究发展。
原文摘要 · Abstract (English)
Language models have recently advanced into the realm of reasoning, yet it is through multimodal reasoning that we can fully unlock the potential to achieve more comprehensive, human-like cognitive capabilities. This survey provides a systematic overview of the recent multimodal reasoning approaches, categorizing them into two levels: language-centric multimodal reasoning and collaborative multimodal reasoning. The former encompasses one-pass visual perception and active visual perception, where vision primarily serves a supporting role in language reasoning. The latter involves action generation and state update within reasoning process, enabling a more dynamic interaction between modalities. Furthermore, we analyze the technical evolution of these methods, discuss their inherent challenges, and introduce key benchmark tasks and evaluation metrics for assessing multimodal reasoning performance. Finally, we provide insights into future research directions from the following two perspectives: (i) from visual-language reasoning to omnimodal reasoning and (ii) from multimodal reasoning to multimodal agents. This survey aims to provide a structured overview that will inspire further advancements in multimodal reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。