让大模型在隐空间自动生成几何辅助线,提升复杂图形推理能力
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
- 在隐空间学习连续视觉表示,无需像素渲染或外部工具
- 三阶段课程训练+强化学习优化,显著提升几何推理准确率
- 专为依赖辅助构造的几何题设计,适合数学推理研究者
尽管多模态推理取得进展,但多模态大语言模型(MLLMs)在表示辅助几何构造方面仍面临根本挑战。这些构造不在原始图中,需在定理应用前引入。现有方法多依赖显式构造范式,如基于文本的几何描述、推理中视觉标记交错、工具增强的几何执行。但这些方法或无法忠实表达复杂空间关系,或导致离散符号与连续几何结构间的表示错配,或依赖外部能力而阻碍端到端优化。为此,我们提出LatentGeo框架,通过学习连续隐空间视觉表示来内化辅助几何构造,无需像素级渲染或外部执行器。设计三阶段课程,逐步通过辅助视觉监督对齐并内化这些隐表示;再引入LaGDPO,一种隐空间感知的强化学习过程,在策略优化中稳定隐表示并提升最终任务正确性。为系统评估构造中心的表示质量,我们构建了GeoAux基准,聚焦视觉依赖的几何问题,并在GeoAux和MathVerse上进行实验。结果表明,LatentGeo在需要辅助构造的几何推理任务中取得显著提升。大量分析与消融实验进一步验证了各组件的有效性。
原文摘要 · Abstract (English)
Despite recent advances in multimodal reasoning, representing auxiliary geometric constructions remains a fundamental challenge for multimodal large language models (MLLMs). Such constructions are absent from the original diagram and must be introduced before theorems apply. Existing approaches predominantly rely on explicit construction paradigms, including text-based geometric specification, visual-token interleaving during reasoning, and tool-augmented geometric execution. However, these methods either fail to faithfully represent complex spatial relationships, incur representation mismatch between discrete symbols and continuous geometric structures, or rely on external capabilities that hinder end-to-end optimization. To address these limitations, we propose LatentGeo, a framework that learns continuous latent visual representations to internalize auxiliary geometric constructions without pixel-level rendering or external executors. We design a three-stage curriculum that progressively aligns and internalizes these latent representations through auxiliary visual supervision, followed by LaGDPO, a latent-aware reinforcement learning procedure that stabilizes latent representations during policy optimization while improving end-task correctness. To systematically evaluate construction-centric representation quality, we introduce GeoAux, a new benchmark targeting visually dependent geometry problems, and conduct experiments on GeoAux and MathVerse. Results show that LatentGeo achieves substantial gains on geometric reasoning tasks, particularly those requiring auxiliary constructions. Extensive analyses and ablation studies further validate the effectiveness of each component in our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。