用比赛式比较提升长文本生成的强化学习奖励效果
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation

- 将相同问题的回答进行多轮比赛,以相对优劣生成奖励
- 在Deep Research Bench上比最强基线高4.52分
- 适合需要高质量长文本生成的研究者
开放域长文本生成中的强化学习面临可靠参考答案和自动评估指标缺失的挑战。现有基于评分标准的方法多依赖逐项LLM评分,但绝对分数难以跨复杂回复校准,对同查询生成物区分度弱,且优化过程中易饱和。本文提出Tournament-GRPO,一种组内比赛式奖励框架,通过多次多轮比赛将基于评分标准的LLM判断转化为相对奖励。该方法在同查询生成物间进行对比,累积比赛结果并归一化为组内奖励用于GRPO训练。在Deep Research Bench上的实验表明,Tournament-GRPO持续优于现有奖励设计基线,整体得分提升4.52分。进一步分析显示,比赛奖励具有良好的效果-效率权衡,且比赛设计影响训练动态。结果表明,基于评分标准的比赛比较可为开放域长文本生成的强化学习提供有效奖励信号。
原文摘要 · Abstract (English)
Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailable. Existing rubric-based methods typically rely on pointwise LLM-as-a-judge scoring, but absolute scores are difficult to calibrate across complex responses, may provide weak discrimination among same-query rollouts, and can become saturated during optimization. We propose Tournament-GRPO, a group-wise reward framework that converts rubric-guided LLM judgments into relative rewards through repeated multi-round tournaments among same-query rollouts. Tournament-GRPO compares candidates within groups, accumulates tournament outcomes, and normalizes them into group-wise rewards for GRPO training. Experiments on Deep Research Bench show that Tournament-GRPO consistently outperforms existing reward-design baselines, achieving a 4.52-point overall-score improvement over the strongest baseline. Further analyses show that tournament rewards provide a favorable effectiveness--efficiency trade-off and that tournament design affects training dynamics. These results suggest that rubric-guided tournament comparison provides an effective reward signal for reinforcement learning in open-ended long-form generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。