arXiv:2603.03202cs.CL2026-03被引 1

用代码代理自动生成更难的数学题,突破难题稀缺瓶颈。

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

  • 设计多智能体框架,通过代码执行探索演化数学题
  • 经充分探索后生成结构新颖且更难的可解新题
  • 适合想提升数学推理能力的模型开发者

随着大语言模型(LLMs)数学能力向国际数学奥林匹克(IMO)和科研水平迈进,高质量、高难度题目的匮乏已成为训练、评估和自演化的关键瓶颈。与此同时,近期代码代理在自主编程与推理方面展现出强大能力,表明代码执行可作为数学实验的可扩展环境。本文探究代码代理自主演化已有数学题为更复杂变体的潜力。我们提出一个多元智能体框架,在生成新题的同时验证其可解性与难度提升。实验表明,经过充分的测试期探索,代码代理能生成结构上与原题不同且更具挑战性的可解问题。本工作提供了实证证据:代码驱动的代理可在可扩展计算环境中有效合成高难度数学推理题目。代码与数据已公开于 https://github.com/TarferSoul/Code2Math。

原文摘要 · Abstract (English)

As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significant bottleneck for training, evaluation and self-evolution of LLMs. Simultaneously, recent code agents have demonstrated sophisticated skills in agentic coding and reasoning, suggesting that code execution can serve as a scalable environment for mathematical experimentation. In this paper, we investigate the potential of code agents to autonomously evolve existing math problems into more complex variations. We introduce a multi-agent framework designed to perform problem evolution while validating the solvability and increased difficulty of the generated problems. Our experiments demonstrate that, given sufficient test-time exploration, code agents can synthesize new, solvable problems that are structurally distinct from and more challenging than the originals. This work provides empirical evidence that code-driven agents can serve as a viable mechanism for synthesizing high-difficulty mathematical reasoning problems within scalable computational environments. Code and data is available at https://github.com/TarferSoul/Code2Math.

数学推理代码代理问题生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。