针对角色扮演对话的奖励建模难题,提出新基准与模型,显著提升叙事连贯性与风格一致性。
RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- 基于连续隐式偏好构建角色扮演奖励模型,实现多策略一致性监督。
- 在七项细粒度能力上,模型平均超越现有方法24%以上。
- 适合需要高保真角色表现的对话系统研究与开发人员。
奖励建模已成为对齐大语言模型与人类偏好的关键。然而在主观且开放的角色扮演对话领域,现有奖励模型性能严重退化,难以捕捉细腻且基于人物设定的人类判断。为此,我们提出了RoleRMBench,首个系统性的角色扮演对话奖励建模基准,涵盖从叙事管理到角色一致性和参与度在内的七项细粒度能力。在RoleRMBench上的评估显示,通用奖励模型与人类判断之间存在巨大且一致的差距,尤其在叙事和风格维度。我们进一步提出RoleRM,一种基于连续隐式偏好(CIP)训练的奖励模型,将主观评价重构为多种结构策略下的连续一致成对监督。全面实验表明,RoleRM在平均性能上超越强开源与闭源奖励模型超过24%,在叙事连贯性和风格保真度方面均有显著提升。研究结果强调了连续偏好表示与标注一致性的重要性,为以人为本的对话系统中的主观对齐奠定了基础。
原文摘要 · Abstract (English)
Reward modeling has become a cornerstone of aligning large language models (LLMs) with human preferences. Yet, when extended to subjective and open-ended domains such as role play, existing reward models exhibit severe degradation, struggling to capture nuanced and persona-grounded human judgments. To address this gap, we introduce RoleRMBench, the first systematic benchmark for reward modeling in role-playing dialogue, covering seven fine-grained capabilities from narrative management to role consistency and engagement. Evaluation on RoleRMBench reveals large and consistent gaps between general-purpose reward models and human judgment, particularly in narrative and stylistic dimensions. We further propose RoleRM, a reward model trained with Continuous Implicit Preferences (CIP), which reformulates subjective evaluation as continuous consistent pairwise supervision under multiple structuring strategies. Comprehensive experiments show that RoleRM surpasses strong open- and closed-source reward models by over 24% on average, demonstrating substantial gains in narrative coherence and stylistic fidelity. Our findings highlight the importance of continuous preference representation and annotation consistency, establishing a foundation for subjective alignment in human-centered dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。