arXiv:2412.11936cs.CL2024-12ACL综述被引 71

首份多模态大模型数学推理综述,梳理200+论文核心进展。

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

  • 系统梳理2021年以来200+研究,聚焦多模态数学推理方法
  • 提出三大维度框架:评测基准、方法体系与关键挑战
  • 揭示5大阻碍通用智能实现的核心难题,指引未来方向

数学推理是人类认知的核心,广泛应用于教育解题与科学发现。随着通用人工智能(AGI)的发展,将大语言模型(LLMs)与数学推理任务结合日益重要。本文首次全面分析多模态大语言模型(MLLMs)时代的数学推理研究。我们综述了2021年以来超过200篇相关论文,重点探讨数学-大语言模型(Math-LLMs)的前沿进展,涵盖多模态场景。从三个维度展开:评测基准、方法体系与挑战。特别分析了多模态数学推理流程,以及(多模态)大语言模型的作用与对应方法。最后,识别出五大制约该领域实现通用智能的关键挑战,并为提升多模态推理能力提供未来研究方向。本综述为研究社区推进大模型复杂多模态推理能力提供了关键资源。

原文摘要 · Abstract (English)

Mathematical reasoning, a core aspect of human cognition, is vital across many domains, from educational problem-solving to scientific advancements. As artificial general intelligence (AGI) progresses, integrating large language models (LLMs) with mathematical reasoning tasks is becoming increasingly significant. This survey provides the first comprehensive analysis of mathematical reasoning in the era of multimodal large language models (MLLMs). We review over 200 studies published since 2021, and examine the state-of-the-art developments in Math-LLMs, with a focus on multimodal settings. We categorize the field into three dimensions: benchmarks, methodologies, and challenges. In particular, we explore multimodal mathematical reasoning pipeline, as well as the role of (M)LLMs and the associated methodologies. Finally, we identify five major challenges hindering the realization of AGI in this domain, offering insights into the future direction for enhancing multimodal reasoning capabilities. This survey serves as a critical resource for the research community in advancing the capabilities of LLMs to tackle complex multimodal reasoning tasks.

数学推理多模态大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。