arXiv:2601.17489cs.LGcs.CL2026-01Conference of the …

让模型看懂几何图并逻辑推理,准确率提升10%。

SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving

论文配图:SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
图 1 · 摘自论文原文
  • 用专门模块提取图形中的空间关系信息
  • 在解题过程中注入空间理解,准确率提升10个百分点
  • 适合需要图形推理的数学竞赛或教育场景

多模态小中型语言模型在融合视觉与文本信息方面表现强劲,但在视觉理解与数学推理,特别是视觉融合度高的几何问题上仍存在显著局限。现有模型难以准确分解复杂视觉输入,并实现感知与结构化推理的有效连接,导致性能不佳。为此,我们提出SpatialMath——一种空间理解增强的符号推理框架,将空间表征融入结构化推理链。该框架采用专用感知模块从视觉图示中提取空间基础表征,捕捉关键几何结构与空间关系,并系统性地将其注入符号推理过程,实现视觉理解驱动的结构化推理。为此,我们构建了MATHVERSE-PLUS数据集,包含视觉密集型数学问题的结构化视觉解析与逐步推理路径。SpatialMath显著优于多个强基线模型,在视觉密集场景下比带数据增强的监督微调高出最高10个百分点。鲁棒性分析表明,强化的空间表征直接提升推理准确性,验证了在多模态小中型语言模型中构建结构化感知-推理流程的必要性。

原文摘要 · Abstract (English)

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning, particularly in geometric problems with diverse levels of visual infusion. Current models struggle to accurately decompose intricate visual inputs and connect perception with structured reasoning, leading to suboptimal performance. To address these challenges, we propose SpatialMath, a novel Spatial Comprehension-Infused Symbolic Reasoning Framework designed to integrate spatial representations into structured symbolic reasoning chains. SpatialMath employs a specialized perception module to extract spatially-grounded representations from visual diagrams, capturing critical geometric structures and spatial relationships. These representations are then methodically infused into symbolic reasoning chains, facilitating visual comprehension-aware structured reasoning. To this end, we introduce MATHVERSE-PLUS, a novel dataset containing structured visual interpretations and step-by-step reasoning paths for vision-intensive mathematical problems. SpatialMath significantly outperforms strong multimodal baselines, achieving up to 10 percentage points improvement over supervised fine-tuning with data augmentation in vision-intensive settings. Robustness analysis reveals that enhanced spatial representations directly improve reasoning accuracy, reinforcing the need for structured perception-to-reasoning pipelines in MSLMs.

数学推理空间理解视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。