arXiv:2605.16371cs.CVcs.AI2026-05被引 1

构建可验证的几何推理数据集,提升模型对图形依赖题的解答能力。

GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning

论文配图:GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
图 1 · 摘自论文原文
  • 用符号化求解器生成精确几何答案,结合渲染管道生成高精度图示。
  • 推出127K带符号真值的问题,使Qwen3-VL模型在几何测试中提升22.21%。
  • 适合需要精准逻辑推理的数学视觉任务研究者使用。

大型多模态模型在几何推理上常因视觉幻觉和缺乏数学精确的思维链(CoT)数据而表现不佳。为此,我们提出GeoSym引擎——一种自动化且可扩展的神经符号框架。通过类型条件语法和解析式SymGT求解器,该框架生成精确的符号真值,并无缝集成于稳健的渲染流水线,产出高精度几何图示。基于此引擎,我们构建了GeoSym127K数据集,包含51,000张高分辨率图像、127,000个带符号真值的问题,以及55,000个经验证的答案-思维链问答对。我们还设计了专家精标511个复杂样本的GeoSym-Bench评估套件。通过大量监督微调(SFT),我们证明GeoSym显著提升了对依赖图示与多步推理的任务表现。Qwen3-VL-8B模型在MathVerse Vision-Only子集上绝对提升22.21%,在WeMath上达到61.52%(+6.19%),缓解长程逻辑断裂问题,优于如Doubao-1.8等闭源先进模型。进一步采用可验证奖励强化学习(RLVR)结合GRPO方法发现:从结构化SFT检查点初始化可大幅提高性能上限,超越零样本强化学习。由确定性精确匹配信号驱动,展现出可验证推理合成的强大扩展潜力。数据集与代码已开源。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address this, we propose the GeoSym Engine, an automated and scalable neuro-symbolic framework. By leveraging a type-conditional grammar and an analytic SymGT Solver, it derives exact symbolic ground truths and seamlessly integrates with a robust rendering pipeline to produce high-precision geometric diagrams. Using this engine, we construct GeoSym127K, a difficulty-stratified dataset featuring 51K high-resolution images, 127K questions with symbolic ground truths, and 55K answer-verified CoT QA pairs. We also introduce GeoSym-Bench, an expert-curated suite of 511 complex samples for rigorous evaluation. Through extensive supervised fine-tuning (SFT), we demonstrate that GeoSym drives concentrated improvements specifically on diagram-dependent and multi-step geometry tasks. Our Qwen3-VL-8B model gains an absolute +22.21% on the MathVerse Vision-Only subset and reaches 61.52% (+6.19% improvement) on WeMath, mitigating long-horizon logic fragmentation and outperforming advanced closed-source models like Doubao-1.8. Furthermore, applying Reinforcement Learning with Verifiable Rewards (RLVR) via GRPO reveals that initializing from structural SFT checkpoints substantially elevates the performance ceiling over zero-shot RL. Driven by deterministic exact-match signals, this showcases the robust scaling potential of our verifiable reasoning synthesis. Datasets and code are available at https://huggingface.co/datasets/Tomie0506/GeoSym127K and https://github.com/Tomie56/GeoSym127K.

几何推理符号化多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。