用对抗演化生成数学题,1000条合成数据胜过千条真实数据。
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis

- 分两步生成:先设计逻辑骨架,再转成自然语言。
- 用1000条合成题训练,性能超过多个真实数据集。
- 适合需要高质量数学推理数据的研究者。
无须人工先验的情况下合成高质量数学推理数据仍是重大挑战。现有方法多依赖种子数据变异或简单提示工程,常出现模式坍缩且逻辑复杂度有限。本文提出一种分层合成框架,将数据合成建模为约束图上的无监督优化问题,随后进行语义实例化,而非直接文本生成。引入立法者-执行者范式:立法者对抗性地演化编码问题约束的结构化生成蓝图,执行者将这些规范实例化为多样化的自然语言场景。这种骨架设计与语言实现的解耦,使重点聚焦于构建复杂多样的逻辑结构,从而引导高质量数据合成。在涵盖Qwen、Llama、Mistral和Gemma系列共10个模型的实验中,使用1000条合成样本微调的模型,在8个数学基准上表现优于规模相当的主流数据集(LIMO、s1K),展现出更强的分布外泛化能力。
原文摘要 · Abstract (English)
Synthesizing high-quality mathematical reasoning data without human priors remains a significant challenge. Current approaches typically rely on seed data mutation or simple prompt engineering, often suffering from mode collapse and limited logical complexity. This paper proposes a hierarchical synthesis framework that formulates data synthesis as an unsupervised optimization problem over a constraint graph followed by semantic instantiation, rather than treating it as a direct text generation task. We introduce a Legislator-Executor paradigm: The Legislator adversarially evolves structured generation blueprints encoding the constraints of the problem, while the Executor instantiates these specifications into diverse natural language scenarios. This decoupling of skeleton design from linguistic realization enables a prioritized focus on constructing complex and diverse logical structures, thereby guiding high-quality data synthesis. Experiments conducted on a total of 10 models across the Qwen, Llama, Mistral, and Gemma series demonstrate that our method achieves notable results: models fine-tuned on 1K synthesized samples outperform widely-used datasets of comparable scale (LIMO, s1K) across eight mathematical benchmarks, exhibiting superior out-of-distribution generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。