arXiv:2509.04903cs.CL2025-09被引 11

用细粒度约束提升长文本生成的强化学习效果

ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning

  • 将指令拆解为可自适应的细粒度约束条件
  • 在WritingBench上比现有方法提升18.63%
  • 适合需要精准控制输出质量的研究与应用

长文本生成已成为大语言模型的重要且具挑战性的应用。现有研究受限于高质量长文本数据稀缺,以及依赖粗粒度通用指标(如连贯性、有用性),忽视了真实场景中任务特有的细微要求。为此,我们提出一种基于自适应约束增强奖励的长文本生成强化学习框架(ACE-RL)。该框架首先将每个指令分解为覆盖长文本生成关键维度的细粒度自适应约束条件;随后设计奖励机制,依据响应对各约束的满足程度量化质量,将主观评价转化为约束验证;最后利用强化学习优化大模型。实验表明,ACE-RL在WritingBench上分别比现有SFT和强化学习基线提升18.63%和7.61%,其最优模型甚至超越GPT-4o达8.76%,为长文本生成提供了更有效的训练范式。

原文摘要 · Abstract (English)

Long-form generation has become a critical and challenging application for Large Language Models (LLMs). Existing studies are limited by their reliance on scarce, high-quality long-form response data and their focus on coarse-grained, general-purpose metrics (e.g., coherence and helpfulness), overlooking the nuanced, scenario-specific requirements of real-world tasks. To address these limitations, we propose a framework utilizing Adaptive Constraint-Enhanced reward for long-form generation Reinforcement Learning (ACE-RL). ACE-RL first decomposes each instruction into a set of fine-grained, adaptive constraint criteria spanning key dimensions of long-form generation tasks. Subsequently, we design a reward mechanism to quantify the response quality based on their satisfaction over corresponding constraints, converting subjective quality evaluation into constraint verification. Finally, we leverage reinforcement learning to optimize LLMs using these fine-grained signals. Experimental results show that ACE-RL significantly outperforms existing SFT and RL baselines by 18.63% and 7.61% on WritingBench, and our top-performing model even surpasses proprietary systems like GPT-4o by 8.76%, providing a more effective training paradigm in long-form generation scenarios.

长文本生成强化学习约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。