arXiv:2505.10218cs.CL2025-05被引 13

用可验证奖励提升大模型角色一致性,让对话更连贯可信。

RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward

  • 通过可验证奖励机制,量化评估角色关键信息匹配度。
  • 14B模型在剧本知识和对话记忆任务上分别达88.04%和88.65%准确率。
  • 适合需要长期角色一致性的虚拟助手、游戏NPC等场景。

角色扮演对话智能体(RPCAs)始终面临角色一致性难以维持的挑战。本文提出RAIDEN-R1,一种融合可验证角色感知奖励(VRAR)的强化学习框架。该方法采用单个与多项式挖掘策略,通过评估角色特异性关键信息生成可量化的奖励信号。同时,通过多大模型协作构建高质量、角色感知的思维链数据集,并用于增强推理连贯性。在RAIDEN基准上的实验表明,14B-GRPO模型在基于剧本的知识和对话记忆任务中分别达到88.04%和88.65%的准确率,优于基线模型且具备强鲁棒性。案例分析显示,模型在处理冲突上下文线索和保持第一人称叙述一致性方面显著提升。本工作填补了角色扮演训练中非可量化性的空白,揭示了角色感知推理模式,推动了角色扮演智能体的发展。

原文摘要 · Abstract (English)

Role-playing conversational agents (RPCAs) face persistent challenges in maintaining role consistency. To address this, we propose RAIDEN-R1, a novel reinforcement learning framework that integrates Verifiable Role-Awareness Reward (VRAR). The method introduces both singular and multi-term mining strategies to generate quantifiable rewards by assessing role-specific keys. Additionally, we construct a high-quality, role-aware Chain-of-Thought dataset through multi-LLM collaboration, and implement experiments to enhance reasoning coherence. Experiments on the RAIDEN benchmark demonstrate RAIDEN-R1's superiority: our 14B-GRPO model achieves 88.04% and 88.65% accuracy on Script-Based Knowledge and Conversation Memory metrics, respectively, outperforming baseline models while maintaining robustness. Case analyses further reveal the model's enhanced ability to resolve conflicting contextual cues and sustain first-person narrative consistency. This work bridges the non-quantifiability gap in RPCA training and provides insights into role-aware reasoning patterns, advancing the development of RPCAs.

角色扮演强化学习大模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。