让AI理解几何题中的自然语言描述并准确定位图形元素
GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- 用结构化几何语言生成海量合成数据训练模型
- 强化学习微调使模型准确率提升显著,超越传统监督学习
- 适合研究多模态数学推理与视觉语言对齐的学者
基于自然语言查询精准定位几何图中的点、形状和空间关系是智能解几何题的关键挑战。本文提出几何参考表达理解(REC)任务,构建了GeoRef基准数据集,涵盖多样且高质量的标注与查询。针对标注数据稀缺问题,利用结构化几何形式语言生成大规模合成训练数据,覆盖广泛几何概念。采用监督微调(SFT)与组相对策略优化(GRPO)两种方法,结果表明GRPO更有效对齐任务奖励,显著提升性能。进一步设计验证-重推机制,通过上下文推理历史修正错误预测,持续提升准确率。即使最先进的多模态大模型在此任务上仍表现不佳,凸显几何对齐能力的重要性。在GeoRef上训练的模型在下游几何推理任务中也取得可测量提升,证明REC是实现多模态数学理解的重要基础。
原文摘要 · Abstract (English)
AI-driven geometric problem solving is a complex vision-language task that requires accurate diagram interpretation, mathematical reasoning, and robust cross-modal grounding. A foundational yet underexplored capability for this task is the ability to identify and interpret geometric elements based on natural language queries. To address this, we introduce the task of Referring Expression Comprehension (REC) for geometric problems, which evaluates whether models can localize points, shapes, and spatial relations in diagrams in response to textual prompts. We present GeoRef, a benchmark dataset constructed from existing geometric problem corpora, featuring diverse, high-quality annotations and queries. Due to the lack of annotated data for this task, we generate a large-scale synthetic training dataset using a structured geometric formal language, enabling broad coverage of geometric concepts and facilitating model adaptation. We explore two fine-tuning approaches: Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). Our results show that GRPO significantly outperforms SFT by better aligning model behavior with task-specific rewards. Furthermore, we propose a verify-and-regenerate mechanism that detects incorrect predictions and re-infers answers using contextual reasoning history, further boosting accuracy. Notably, even state-of-the-art Multimodal Large Language Models (MLLMs) struggle with this task, underscoring the necessity of explicitly evaluating and strengthening geometric grounding as a prerequisite for robust geometric problem solving. Moreover, models trained on GeoRef demonstrate measurable improvements on downstream geometric reasoning tasks, highlighting the broader value of REC as a foundation for multimodal mathematical understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。