arXiv:2605.25384cs.CL2026-05

用代码表示几何解题中间步骤,提升模型推理能力

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving

论文配图:GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
图 1 · 摘自论文原文
  • 用程序代码作为几何推理的中间可视化输出
  • 推理与代码生成在隐空间可分离,微调后推理更结构化
  • 代码结构蕴含更多数学符号信息,适合研究模型思维

数学推理是人类智能的标志,需逻辑推导、符号操作与抽象思考。近期多模态大语言模型通过多步推理在几何问题上表现优异。为更好模拟人类解题,中间步骤可引入辅助视觉构造(如新增线段或点),提升几何理解与教学清晰度。本文提出GeoMathCode,以程序化表示作为中间视觉输出。深入分析显示,推理与代码生成可在隐空间解耦,监督微调使推理流形更结构化、信息更丰富。此外,分层语法代码结构在隐空间中形成独立子空间,包含比视觉表征更多的数学符号信息。

原文摘要 · Abstract (English)

Mathematical reasoning is a hallmark of human intelligence, requiring logical deduction, symbolic manipulation, and abstract thinking. Recent multimodal large language models (MLLMs) have demonstrated strong performance on geometry problems through multi-step reasoning. To better emulate human problem-solving, intermediate steps can incorporate auxiliary visual constructions, such as additional lines or points, which improve geometric interpretation and educational clarity. In this work, we introduce the GeoMathCode, where programmatic representations serve as intermediate visual outputs. We further conduct an in-depth analysis of the underlying reasoning geometry. Experimental results show that reasoning and code generation steps can be disentangled in the latent space, while supervised fine-tuning (SFT) makes the reasoning manifold more structured and informative. Moreover, hierarchical syntactic code structures emerge as disentangled latent subspaces, and contain more mathematical symbolic information than visual representations.

几何推理代码表示多模态模型隐空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。