arXiv:2512.24297cs.CL2025-12

让AI边想边画图,提升复杂推理能力

Figure It Out: Improve the Frontier of Reasoning with Executable Visual States

  • 用可执行代码动态生成图形辅助推理
  • 在AIME 2025上比纯文本模型高13.12%
  • 适合需要空间几何推理的数学竞赛题

复杂推理问题常涉及文本未明确定义的空间与几何关系。尽管近期推理模型在多个领域表现良好,但纯文本推理难以捕捉复杂场景中的结构约束。本文提出FIGR,通过端到端强化学习将可执行视觉构建融入多轮推理过程。不同于仅依赖文本思维链,FIGR在推理循环中生成可执行代码以构建图形,外部化中间假设。自适应奖励机制智能控制视觉构建时机,使模型更一致地推理难以从文本推断的全局隐含属性。在八个挑战性数学基准测试中,FIGR显著优于强文本基线,在AIME 2025上提升13.12%,在BeyondAIME上提升11.00%。结果表明,精准可控的图形构造能有效增强复杂推理能力。

原文摘要 · Abstract (English)

Complex reasoning problems often involve implicit spatial and geometric relationships that are not explicitly encoded in text. While recent reasoning models perform well across many domains, purely text-based reasoning struggles to capture structural constraints in complex settings. In this paper, we introduce FIGR, which integrates executable visual construction into multi-turn reasoning via end-to-end reinforcement learning. Rather than relying solely on textual chains of thought, FIGR externalizes intermediate hypotheses by generating executable code that constructs diagrams within the reasoning loop. An adaptive reward mechanism selectively regulates when visual construction is invoked, enabling more consistent reasoning over latent global properties that are difficult to infer from text alone. Experiments on eight challenging mathematical benchmarks demonstrate that FIGR outperforms strong text-only chain-of-thought baselines, improving the base model by 13.12% on AIME 2025 and 11.00% on BeyondAIME. These results highlight the effectiveness of precise, controllable figure construction of FIGR in enhancing complex reasoning ability.

空间推理代码生成数学竞赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。