arXiv:2506.05760cs.CL2025-06ACL被引 10

通过自适应课程强化学习,提升大模型长文本生成能力。

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

  • 基于学习潜力选择数据,动态调整任务难度
  • 无验证奖励下仍能提供有效训练信号
  • 70亿参数模型效果优于监督微调,且可迁移到长推理

大语言模型在长文本生成方面已取得显著进展,但现有训练方法仍有局限:监督微调受数据饱和和性能瓶颈制约,而基于可验证奖励的强化学习虽在数学与代码等可验证领域成功应用,却难以直接迁移至开放式长文本生成,因缺乏真实答案。为此,我们提出Writing-RL:一种自适应课程强化学习框架,以突破监督微调的上限。该框架包含三个核心组件:基于置信度边缘的数据筛选策略,优先选择高学习潜力样本;成对比较奖励机制,在无明确标准时提供判别性学习信号;动态参考调度方法,根据模型演进自动调节任务难度。在70亿参数写作者模型上的实验表明,Writing-RL显著优于强基线监督微调。此外,经长输出强化学习训练的模型在长输入推理任务上展现出意外的良好泛化能力,为重新思考长上下文训练提供了新思路。

原文摘要 · Abstract (English)

Recent advances in Large Language Models(LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited: Supervised Fine-Tuning (SFT) remains constrained by data saturation and performance ceilings, while Reinforcement Learning with Verifiable Reward (RLVR), though successful in verifiable domains like math and code, cannot be directly migrated to open-ended long-form writing due to a lack of ground-truths. To further advance long-form writing, we present Writing-RL: an Adaptive Curriculum Reinforcement Learning framework to advance long-form writing capabilities beyond SFT. The framework consists of three key components: Margin-aware Data Selection strategy that prioritizes samples with high learning potential, Pairwise Comparison Reward mechanism that provides discriminative learning signals in the absence of verifiable rewards, and Dynamic Reference Scheduling approach, which plays a critical role by adaptively adjusting task difficulty based on evolving model performance. Experiments on 7B-scale writer models show that Writing-RL effectively improves long-form writing performance over strong SFT baselines. Furthermore, we observe that models trained with long-output RL generalize surprisingly well to long-input reasoning tasks, potentially offering a promising perspective for rethinking long-context training.

长文本生成强化学习自适应课程大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。