通过动态筛选中等难度任务,提升大模型推理强化学习的训练效率
Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning

- 基于任务难度的在线筛选机制,选择中间难度样本提升学习效率
- 在多个数学推理数据集上实现速度提升超50%、性能最高+12%的增益
- 适用于需要高效训练大模型推理能力的研究者与工程师
最近基于可验证奖励的强化学习(RLVR)进展表明,大型语言模型在接收可验证信号后能显著提升推理能力。然而由于奖励稀疏,模型表现高度依赖于样本难度的选择。本文首次对在线难度感知筛选进行形式化分析,证明期望策略改进下界由任务成功率方差决定,因此选择中等难度任务可最大化学习效率。进一步证明平衡筛选能最大化该下界,带来更优性能与更高样本效率。在多个数学推理基准上的实验验证,平衡筛选持续加速收敛并提升最终性能,相比标准GRPO方法在不足一半训练步数内实现最高+12%的提升。通过将分析扩展至多种奖励分布,本文为未来RLVR课程设计提供理论基础,经理论推导与大规模实证结果共同验证。
原文摘要 · Abstract (English)
Recent advances in reinforcement learning with verifiable rewards (RLVR) show that large language models enhance their reasoning abilities when trained with verifiable signals. However, due to reward sparsity, effectiveness depends heavily on selecting samples of appropriate difficulty. In this work, we present a formal analysis of online difficulty-aware filtering and establish its theoretical foundations. We show that expected policy improvement is lower-bounded by the variance of task-level success probabilities, implying that selecting tasks of intermediate difficulty maximizes learning efficiency. Building on this, we demonstrate that balanced filtering maximizes this lower bound, leading to superior performance and sample efficiency. Evaluations across multiple math reasoning benchmarks validate that balanced filtering consistently enhances convergence speed and final performance, achieving up to +12% gains in less than half the training steps of standard GRPO. By extending our analysis to various reward distributions, we provide a principled foundation for future RLVR curriculum strategies, confirmed through both theoretical analysis and extensive empirical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。