arXiv:2605.20743cs.CVcs.CL2026-05

让AI画图时实时验证几何约束,提升解题准确率

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction

论文配图:Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction
图 1 · 摘自论文原文
  • 用GeoGebra引擎交互式构建可验证的几何图
  • 在几何问题上达到95.9%的构造通过率
  • 适合需要精确几何推理的研究者和教育应用

视觉语言模型在解决几何问题上的准确率持续提升,但其中间推理过程隐含且不可验证:文本或绘图代码表达的关系无法保证对应满足约束的几何构型。现有基于像素渲染或单次脚本的外部化方法无法提供每一步的精确几何保障。通过代数定义强制几何关系,使工作区成为受约束检查的动态画布。我们提出Draw2Think框架,将几何推理从隐式空间推演转变为与GeoGebra约束引擎的代理式交互。在提议-绘制-验证循环中,模型将假设外化到可执行画布,测量精确几何量,并将结构化观察反馈给模型,使后续推理基于经验证的共享状态。该机制使两个属性可独立审计:模型级构造保真度(画布是否实现预期构型)与引擎级测量忠实度(画布约束下的精确值与关系)。在构造、结果和渲染评估中,Draw2Think在GeoGoal上通过95.9%的谓词级与84.0%的严格问题级构造检查,在平面与立体基准上结果准确率提升最高达4.1%/16.4%,在GenExam-math上取得68.2%/90.5%的严格/宽松渲染得分。

原文摘要 · Abstract (English)

Vision-language models solve geometry problems with rising accuracy, yet their intermediate states remain latent and unverifiable: a relation expressed in textual reasoning or drawing code carries no guarantee that a constraint-satisfying configuration realizes it. We observe that existing externalization methods based on rendered pixels or one-shot scripts fail to provide exact, per-action geometric guarantees. Enforcing geometric relations by algebraic definition closes this gap: the workspace becomes a constraint-checked evolving canvas. We present Draw2Think, a framework that recasts geometric reasoning from latent spatial inference into agentic interaction with the GeoGebra constraint engine. In a Propose-Draw-Verify loop, Draw2Think externalizes hypotheses onto an executable canvas, measures exact geometric quantities, and feeds structured observations back to the model, so subsequent reasoning proceeds from checked canvas state grounded by the shared workspace. This externalization makes two properties separately auditable: model-level Construction Fidelity (whether the canvas realizes the intended configuration) and engine-level Measurement Faithfulness (exact values and relations from canvas constraints). Across construction, outcome, and rendering evaluations, Draw2Think builds canvases that pass 95.9% predicate-level and 84.0% strict problem-level construction checks on GeoGoal, improves outcome accuracy by up to 4.1%/16.4% on planar/solid benchmarks, and attains 68.2%/90.5% strict/relaxed rendering scores on GenExam-math. Project page is available at https://draw2think.github.io/

几何推理约束引擎AI作图可验证性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。