arXiv:2504.03151cs.CLcs.LG2025-04综述被引 23

梳理多模态推理的核心挑战与优化方法,助力大模型跨模态理解

Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)

  • 对比文本与多模态场景下的推理机制差异
  • 提出后训练优化与推理时策略的实用方案
  • 适合关注大模型跨模态能力的研究者参考

推理是人类智能的核心,支持在多样化任务中进行结构化问题求解。近年来,大语言模型(LLMs)在算术、常识和符号领域显著提升了推理能力。然而,将这些能力有效扩展到多模态场景——即模型需整合视觉与文本输入——仍是重大挑战。多模态推理引入了模态间信息冲突等复杂性,要求模型采用高级解释策略。解决这些问题不仅需要复杂算法,还需稳健的评估方法以衡量推理准确性和连贯性。本文系统综述了文本与多模态大模型中的推理技术,通过全面且最新的对比,明确提炼出核心挑战与机遇,重点突出后训练优化与测试时推理的实用方法。工作为理论框架与实际应用之间搭建桥梁,为未来研究指明方向。

原文摘要 · Abstract (English)

Reasoning is central to human intelligence, enabling structured problem-solving across diverse tasks. Recent advances in large language models (LLMs) have greatly enhanced their reasoning abilities in arithmetic, commonsense, and symbolic domains. However, effectively extending these capabilities into multimodal contexts-where models must integrate both visual and textual inputs-continues to be a significant challenge. Multimodal reasoning introduces complexities, such as handling conflicting information across modalities, which require models to adopt advanced interpretative strategies. Addressing these challenges involves not only sophisticated algorithms but also robust methodologies for evaluating reasoning accuracy and coherence. This paper offers a concise yet insightful overview of reasoning techniques in both textual and multimodal LLMs. Through a thorough and up-to-date comparison, we clearly formulate core reasoning challenges and opportunities, highlighting practical methods for post-training optimization and test-time inference. Our work provides valuable insights and guidance, bridging theoretical frameworks and practical implementations, and sets clear directions for future research.

多模态推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。