arXiv:2509.09013cs.CLcs.AI2025-09EMNLP被引 3

测试视觉语言模型解图像数学题的能力,发现计数是主要瓶颈。

Can Vision-Language Models Solve Visual Math Equations?

论文配图:Can Vision-Language Models Solve Visual Math Equations?
图 1 · 摘自论文原文
  • 将图像方程分解为计数和识别两步,发现计数能力弱于识别。
  • 模型在复杂方程上表现下降,符号推理成新瓶颈。
  • 适合关注多模态推理缺陷的研究者阅读。

尽管在视觉理解与基于语言的推理方面表现强劲,视觉语言模型(VLMs)在需要整合感知与符号计算的任务中仍表现不佳。本文通过图像方程求解任务研究这一局限性:数学方程嵌入图像,变量由物体图标表示,系数需通过计数推断。虽然VLMs在文本方程上表现良好,但在视觉接地的方程上失败。我们分解任务为系数计数与变量识别,发现计数是主要瓶颈,即使识别准确也难以突破。同时,识别与推理组合引入额外错误,凸显多步视觉推理的挑战。随着方程复杂度增加,符号推理本身也成为限制因素。这些发现揭示了当前VLMs在视觉化数学推理中的关键弱点,并指明未来改进方向。

原文摘要 · Abstract (English)

Despite strong performance in visual understanding and language-based reasoning, Vision-Language Models (VLMs) struggle with tasks requiring integrated perception and symbolic computation. We study this limitation through visual equation solving, where mathematical equations are embedded in images, variables are represented by object icons, and coefficients must be inferred by counting. While VLMs perform well on textual equations, they fail on visually grounded counterparts. To understand this gap, we decompose the task into coefficient counting and variable recognition, and find that counting is the primary bottleneck, even when recognition is accurate. We also observe that composing recognition and reasoning introduces additional errors, highlighting challenges in multi-step visual reasoning. Finally, as equation complexity increases, symbolic reasoning itself becomes a limiting factor. These findings reveal key weaknesses in current VLMs and point toward future improvements in visually grounded mathematical reasoning.

视觉语言模型数学推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。