arXiv:2509.21896cs.AI2025-09被引 1

构建百万级几何推理数据集,让AI学会看图解题

GenesisGeo: Technical Report

  • 用合成数据训练模型同时理解图形和符号推理
  • 在奥数几何题上解出29/30道,达到金牌水平
  • 适合对视觉推理、数学AI感兴趣的开发者

近期神经符号几何定理证明系统在欧几里得几何问题上取得进展,通过结合神经引导与符号验证。然而,现有系统几乎仅在符号空间中运行,未充分利用图形提供的直观启发。人类解题时常依赖图形发现非平凡辅助线构造。与此同时,视觉语言模型(VLM)在几何任务上表现不佳,因缺乏高质量含几何图示与推理监督的数据。本文提出GenesisGeo-1M,一个大规模合成的多模态几何推理数据集,包含100万组图文几何问题及其可机器验证的证明轨迹。基于此数据集,我们提出一种多任务训练范式,联合优化文本生成与图示引导的证明生成,促使模型学习视觉关联与符号推导。大量实验表明,我们的GenesisGeo-2B模型在奥数几何基准上达到金牌级表现:在IMO-30上解决29/30题,在IMO-95上解决63/95题,在HAGeo-409上解决278/409题。

原文摘要 · Abstract (English)

Recent neuro-symbolic geometry theorem provers have made significant progress on Euclidean problems by coupling neural guidance with symbolic verification. However, most existing systems operate almost exclusively in a symbolic space, leaving diagram-based intuition largely unused during reasoning. For humans, geometric diagrams provide essential heuristics for identifying non-trivial auxiliary constructions. Meanwhile, visual language models (VLMs) still struggle with geometry due to the lack of high-quality data with geometric diagrams and reasoning supervision. In this paper, we introduce GenesisGeo-1M, a large-scale synthetic dataset for visual geometric reasoning that contains 1M multimodal geometry problems paired with machine-checkable proof traces. Building on this dataset, we formulate geometric learning as a multi-task training paradigm that jointly optimizes text-based proof generation and diagram-grounded proof generation, encouraging models to learn visual grounding and symbolic deduction. Extensive experiments show that our GenesisGeo-2B model achieves gold-medal-level performance on Olympiad geometry benchmarks, solving 29/30 problems on IMO-30, 63/95 on IMO-95, and 278/409 on HAGeo-409.

几何推理视觉语言模型多模态训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。