测试大模型对相似图表的数学推理能力,发现其易受位置误导。
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- 构建1800道中小学数学题,答案图仅细微差异
- 模型准确率随图像相似度升高而显著下降
- 提出对齐策略提升多图文本理解,适合教育与视觉推理研究者
大型多模态模型在视觉与语言融合方面取得显著进展,但在处理多个视觉相似输入时的细粒度比较推理能力仍不足。这类推理在数学与教育场景中至关重要,学习者需区分几乎相同的图表以找出正确解法。为此,我们提出VisioMath,一个包含1800道高质量中小学数学题的基准数据集,所有候选答案均为视觉上极为相似的图表。对主流闭源与开源LMM进行全面评估发现,随着图像间相似度增加,模型准确率持续下降。分析表明,主要失败原因是图像-文本错位:模型未基于文本线索进行推理,而是依赖浅层位置启发式,导致系统性错误。我们进一步探索了三种面向对齐的策略,涵盖训练无关方法与微调,显著提升准确率。希望VisioMath能成为推动多模态模型迈向更深层次图表理解、精确比较推理和多图文本融合的严格基准与催化剂。
原文摘要 · Abstract (English)
Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. However, their capacity to reason over multiple, visually similar inputs remains insufficiently explored. Such fine-grained comparative reasoning is central to real-world tasks, especially in mathematics and education, where learners must often distinguish between nearly identical diagrams to identify correct solutions. To address this gap, we present VisioMath, a curated benchmark of 1,800 high-quality K-12 mathematics problems in which all candidate answers are diagrams with subtle visual similarities. A comprehensive evaluation of state-of-the-art LMMs, covering both leading closed-source systems and widely adopted open-source models, reveals a consistent decline in accuracy as inter-image similarity increases. Analysis indicates that the dominant failure mode stems from image-text misalignment: rather than grounding reasoning in textual cues, models often resort to shallow positional heuristics, resulting in systematic errors. We further explore three alignment-oriented strategies, spanning training-free approaches and finetuning, and achieve substantial accuracy gains. We hope that VisioMath will serve as a rigorous benchmark and catalyst for developing LMMs toward deeper diagram understanding, precise comparative reasoning, and grounded multi-image-text integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。