arXiv:2505.04921cs.CVcs.CL2025-05综述被引 95

综述大模型如何融合多模态信息进行深度推理与规划。

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

  • 按感知-理解-推理-规划四阶段梳理多模态大模型发展路径
  • 提出原生多模态推理模型(N-LMRMs)概念,支持自主决策与规划
  • 适合关注AI通用智能、多模态系统设计的研究者参考

推理是智能的核心,决定系统能否做出决策、得出结论并跨领域泛化。在人工智能中,随着系统越来越多地运行于开放、不确定且多模态的环境,推理能力成为实现鲁棒和自适应行为的关键。大型多模态推理模型(LMRMs)应运而生,整合文本、图像、音频、视频等模态,支持复杂推理能力,旨在实现全面感知、精准理解与深层推理。研究进展使多模态推理从早期任务特定模块化的感知驱动流程,演变为统一的语言中心框架,提升了跨模态理解的一致性。尽管指令微调与强化学习提升了模型推理能力,但全模态泛化、推理深度及代理行为仍面临挑战。本文提出一个结构化综述,围绕四阶段发展路线图,系统梳理该领域的范式转变与新兴能力:第一阶段回顾基于任务特定模块的早期探索,推理隐含于表征、对齐与融合各阶段;第二阶段分析将推理统一至多模态大语言模型的最新方法,如多模态思维链(MCoT)与多模态强化学习,支持更丰富、结构化的推理链条;最后,结合OpenAI O3与O4-mini等基准测试与实验案例,探讨原生多模态推理模型(N-LMRMs)的未来方向,其目标是在复杂真实环境中实现可扩展、自主、适应性的推理与规划。

原文摘要 · Abstract (English)

Reasoning lies at the heart of intelligence, shaping the ability to make decisions, draw conclusions, and generalize across domains. In artificial intelligence, as systems increasingly operate in open, uncertain, and multimodal environments, reasoning becomes essential for enabling robust and adaptive behavior. Large Multimodal Reasoning Models (LMRMs) have emerged as a promising paradigm, integrating modalities such as text, images, audio, and video to support complex reasoning capabilities and aiming to achieve comprehensive perception, precise understanding, and deep reasoning. As research advances, multimodal reasoning has rapidly evolved from modular, perception-driven pipelines to unified, language-centric frameworks that offer more coherent cross-modal understanding. While instruction tuning and reinforcement learning have improved model reasoning, significant challenges remain in omni-modal generalization, reasoning depth, and agentic behavior. To address these issues, we present a comprehensive and structured survey of multimodal reasoning research, organized around a four-stage developmental roadmap that reflects the field's shifting design philosophies and emerging capabilities. First, we review early efforts based on task-specific modules, where reasoning was implicitly embedded across stages of representation, alignment, and fusion. Next, we examine recent approaches that unify reasoning into multimodal LLMs, with advances such as Multimodal Chain-of-Thought (MCoT) and multimodal reinforcement learning enabling richer and more structured reasoning chains. Finally, drawing on empirical insights from challenging benchmarks and experimental cases of OpenAI O3 and O4-mini, we discuss the conceptual direction of native large multimodal reasoning models (N-LMRMs), which aim to support scalable, agentic, and adaptive reasoning and planning in complex, real-world environments.

多模态推理大模型智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。