让AI自己生成合适难度的题目,持续提升推理能力。
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning

- 根据模型当前水平自动筛选中等难度题目作为锚点
- 用少于2000个真实数学题就超越现有方法
- 适合需要高效训练推理模型的研究者
强化学习在提升大语言模型推理能力方面展现出潜力。然而,有效的强化学习训练依赖中等难度样本,面临两大挑战:数据稀缺与难度动态变化——中等难度样本稀少,且随着模型进步逐渐变得简单。现有方法通过生成训练样本缓解稀缺性,但存在无锚点生成、忽略协同进化和难度不匹配等问题。为此,我们提出D²Evo,一种双难度感知的自进化强化学习框架。每轮迭代中,该方法基于当前求解器能力挖掘中等难度锚点,训练提问器生成多样且难度适中的问题,并联合优化两者以实现推理能力的持续提升。大量实验表明,D²Evo在数学推理基准上仅使用少于2000个真实数学样本即优于现有方法,并在通用推理基准上表现出强泛化能力。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has demonstrated potential for enhancing reasoning in large language models (LLMs). However, effective RL training, which requires medium-difficulty training samples, faces two fundamental challenges: Effective Data Scarcity and Dynamic Difficulty Shifts, where medium-difficulty samples are scarce and become trivial as models improve. Existing methods mitigate this scarcity to some extent by generating training samples. However, these approaches suffer from anchor-free generation, ignoring co-evolution, and difficulty mismatch. To address these issues, we propose D$^2$Evo, a Dual Difficulty-aware self-Evolution RL framework. In each iteration, our method mines medium-difficulty anchors based on the current Solver's capability, trains the Questioner to generate diverse questions at appropriate difficulty levels, and jointly optimizes both components to enable progressive reasoning gains. Extensive experiments demonstrate that D$^2$Evo outperforms existing methods on mathematical reasoning benchmarks with fewer than 2K real mathematical samples, and exhibits strong generalization on general reasoning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。