用对比评估提升角色扮演对话的奖励一致性
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
- 通过成对比较取代单样本打分,减少主观偏差
- 在多个评测集上显著提升对话质量与稳定性
- 适合研究角色扮演、对话生成与奖励建模的学者
强化学习微调(RLFT)在可客观验证的任务(如代码生成、数学推理)中表现优异,但在开放式的主观任务(如角色扮演对话)中表现不佳。传统奖励建模依赖独立样本评分,面临评价标准主观和奖励信号不稳两大挑战。受人类评价中显性标准与隐性比较并存的启发,我们提出比较策略优化(CPO),将奖励评估范式从样本级打分转向群体级对比打分。基于同一原则,我们构建了CharacterArena评估框架,包含两个阶段:(1) 上下文感知的多轮角色扮演模拟,(2) 轨迹级对比评估。通过客观轨迹对比实现主观评分,有效降低上下文偏差,提升评估的鲁棒性与公平性。在CharacterEval、CharacterBench和CharacterArena上的实证结果表明,CPO能有效缓解奖励模糊性,显著提升对话质量。
原文摘要 · Abstract (English)
Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles with open-ended subjective tasks like role-playing dialogue. Traditional reward modeling approaches, which rely on independent sample-wise scoring, face dual challenges: subjective evaluation criteria and unstable reward signals.Motivated by the insight that human evaluation inherently combines explicit criteria with implicit comparative judgments, we propose Comparative Policy Optimization (CPO). CPO redefines the reward evaluation paradigm by shifting from sample-wise scoring to comparative group-wise scoring.Building on the same principle, we introduce the CharacterArena evaluation framework, which comprises two stages:(1) Contextualized Multi-turn Role-playing Simulation, and (2) Trajectory-level Comparative Evaluation. By operationalizing subjective scoring via objective trajectory comparisons, CharacterArena minimizes contextual bias and enables more robust and fair performance evaluation. Empirical results on CharacterEval, CharacterBench, and CharacterArena confirm that CPO effectively mitigates reward ambiguity and leads to substantial improvements in dialogue quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。