arXiv:2601.01931cs.LGcs.AI2026-01

让模型自己生成并优化数学题,提升推理能力。

DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems

  • 模型自动生成新题目,动态调整难度适应训练进度。
  • 相比固定数据集,强化学习训练效率提升明显。
  • 适合研究模型泛化与自动命题系统的人参考。

近期推理模型在数学和编程任务上表现优异,但多数方法依赖静态数据集,易导致记忆现象并限制泛化能力。我们提出DéjàQ框架,通过联合演化多样化的合成数学问题与模型训练,使问题随模型能力动态调整,优化可学性。提出两种由大语言模型驱动的变异策略:一是修改上下文细节,二是直接改变问题结构。实验表明,模型能生成新颖且有意义的问题,且这些自变异数据显著提升强化学习训练效果。我们分析了生成问题的有效性及计算开销。结果表明,动态演化训练数据可有效增强数学推理能力,具有广泛适用性,代码将开源。

原文摘要 · Abstract (English)

Recent advances in reasoning models have yielded impressive results in mathematics and coding. However, most approaches rely on static datasets, which have been suggested to encourage memorisation and limit generalisation. We introduce DéjàQ, a framework that departs from this paradigm by jointly evolving a diverse set of synthetic mathematical problems alongside model training. This evolutionary process adapts to the model's ability throughout training, optimising problems for learnability. We propose two LLM-driven mutation strategies in which the model itself mutates the training data, either by altering contextual details or by directly modifying problem structure. We find that the model can generate novel and meaningful problems, and that these LLM-driven mutations improve RL training. We analyse key aspects of DéjàQ, including the validity of generated problems and computational overhead. Our results underscore the potential of dynamically evolving training data to enhance mathematical reasoning and indicate broader applicability, which we will support by open-sourcing our code.

数学推理自生成数据强化学习演化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。