无需标量奖励,通过对比偏好与分组多样性提升开放生成质量
Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation

- 用成对偏好替代标量奖励,减少人工标注成本
- 引入分组多样性奖励,使生成内容更丰富不重复
- 适合需要高质量多样性的开放域生成任务
当前强化学习方法在可验证场景中表现良好,但开放生成任务难以验证输出正确性,训练奖励模型成本高昂。且传统强化学习易导致多样性崩溃,产生刻板输出。本文提出PPR-GDE方法,无需标量奖励,通过成对偏好保留主观评价结构,利用响应顺序轮换缓解评判者偏见,并引入分组多样性奖励,显式促进一组回复间的语义分散。所有奖励信号整合为统一的组相对策略优化目标。在角色扮演任务上的实验表明,PPR-GDE在对齐质量和表达多样性上均优于强基线。分析显示,成对偏好对主观对齐至关重要,而多样性度量是实现优异表达多样性和广泛语义覆盖的关键。
原文摘要 · Abstract (English)
Current reinforcement learning(RL) methods are broadly applicable and powerful in verifiable settings where scalar rewards can be provided. However, in open-ended generation tasks, verifying the correctness of responses remains challenging, and training reward models incurs substantial computational and annotation costs. Moreover, reinforcement learning (RLVR) often leads to diversity collapse and produces stereotypical or rigid outputs, outcomes that are particularly undesirable in open-domain scenarios. We propose Pairwise Preference Reward and Group-based Diversity Enhancement (PPR-GDE), a RL method that is more suitable for open-ended generation. PPR-GDE does not require scalar rewards and incorporates group-level diversity into the reward signal, it preserves the comparative structure of subjective evaluation through a pairwise preference reward, mitigates judge position bias via repeated comparisons with swapped response order, and introduces a group-based diversity reward that explicitly encourages semantic dispersion within a response group, all of these reward signals are integrated into a unified group-relative policy optimization objective. We instantiate PPR-GDE on role-playing task, experiments show that PPR-GDE achieves a better alignment quality as well as expressive diversity than strong RL baselines. Further analysis shows that pairwise preference is critical for preference alignment in subjective perspective, while the diversity metric plays an essential role in achieving superior expressive diversity and broader semantic coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。