arXiv:2603.08291cs.AI2026-03ACL综述

系统梳理多模态数学推理的感知、对齐与推理方法

A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning

论文配图:A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
图 1 · 摘自论文原文
  • 从图文中提取数学信息并建立显式对齐
  • 通过可验证推理框架提升中间步骤正确性
  • 适合关注多模态推理评估与未来方向的研究者

多模态数学推理(MMR)近年来受到广泛关注,因其能处理包含文本与视觉信息的数学问题。然而,现有模型在真实场景的视觉数学任务中仍面临显著挑战,常误读图表、无法对齐数学符号与视觉证据,或产生不一致的推理步骤。此外,现有评估主要关注最终答案,而忽视中间步骤的正确性与可执行性。近期研究通过在统一框架中整合结构化感知、显式对齐与可验证推理,逐步解决这些问题。为清晰呈现不同MMR方法的演进路径,本文围绕四个核心问题展开系统综述:(1) 如何从多模态输入中提取信息;(2) 如何表示并对齐文本与视觉内容;(3) 如何实现推理;(4) 如何评估推理过程的整体正确性。最后,讨论开放挑战并展望未来研究方向。

原文摘要 · Abstract (English)

Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both textual and visual modalities. However, current models still face significant challenges in real-world visual math tasks, often misinterpreting diagrams, failing to align mathematical symbols with visual evidence, or producing inconsistent reasoning steps. Moreover, existing evaluations mainly focus on checking final answers rather than verifying the correctness or executability of each intermediate step. A growing body of recent research addresses these issues by integrating structured perception, explicit alignment, and verifiable reasoning within unified frameworks. To establish a clear roadmap for understanding and comparing different MMR approaches, we systematically review them around four fundamental questions: (1) What to extract from multimodal inputs, (2) How to represent and align textual and visual information, (3) How to perform the reasoning, and (4) How to evaluate the correctness of the overall reasoning process. Finally, we discuss open challenges and share our thoughts on future research directions.

多模态数学推理评估综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。