用智能选题系统提升大模型强化学习训练效率与效果。
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
- 通过动态选题机制自动筛选训练问题,优化策略改进。
- 在多个推理任务上提升30%以上,最快提速80%。
- 适合需要高效训练大模型的科研与工程团队。
使用强化学习对大型基础模型进行后训练,通常依赖海量且异构的数据集,因此有效的课程学习既关键又具挑战性。本文提出一种可扩展且完全自动化的强化学习后训练课程学习框架——ACTOR-CURATOR。该框架学习一个神经化课程管理者,通过直接优化预期策略性能提升,从大规模问题库中动态选择训练问题。我们将问题选择建模为非平稳随机老虎机问题,基于在线随机镜面下降推导出一个合理损失函数,并在部分反馈条件下建立后悔率保证。实验表明,ACTOR-CURATOR 在多种具有挑战性的推理基准测试中均优于均匀采样和强基线方法,展现出更优的训练稳定性和效率。特别地,在 AIME2024 上相对最优基线提升 28.6%,在 ARC-1D 上提升 30.5%,最高实现 80% 的加速。结果表明,ACTOR-CURATOR 是一种强大且实用的大规模语言模型后训练方法。
原文摘要 · Abstract (English)
Post-training large foundation models with reinforcement learning typically relies on massive and heterogeneous datasets, making effective curriculum learning both critical and challenging. In this work, we propose ACTOR-CURATOR, a scalable and fully automated curriculum learning framework for reinforcement learning post-training of large language models (LLMs). ACTOR-CURATOR learns a neural curator that dynamically selects training problems from large problem banks by directly optimizing for expected policy performance improvement. We formulate problem selection as a non-stationary stochastic bandit problem, derive a principled loss function based on online stochastic mirror descent, and establish regret guarantees under partial feedback. Empirically, ACTOR-CURATOR consistently outperforms uniform sampling and strong curriculum baselines across a wide range of challenging reasoning benchmarks, demonstrating improved training stability and efficiency. Notably, it achieves relative gains of 28.6% on AIME2024 and 30.5% on ARC-1D over the strongest baseline and up to 80% speedup. These results suggest that ACTOR-CURATOR is a powerful and practical approach for scalable LLM post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。