arXiv:2601.04611cs.CL2026-01被引 6

让角色扮演机器人更懂角色,减少出戏行为。

Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR

  • 通过10个角色特质标签构建认知结构
  • 引用参考回答提升探索与表现
  • 按角色类型动态调整奖励,适配多样角色

当前角色扮演智能体多依赖表面行为模仿,缺乏内在认知一致性,复杂情境下常出现不符合角色的错误。为此,我们提出Character-R1框架,提供全面可验证的奖励信号以实现有效角色感知推理。该框架包含三个核心设计:(1) 认知焦点奖励,强制基于10个角色元素(如世界观)的显式标签分析,以结构化内部认知;(2) 参考引导奖励,利用重叠度指标与参考响应作为优化锚点,增强探索与性能;(3) 角色条件奖励归一化,根据角色类别调整奖励分布,确保在异质角色间稳健优化。大量实验表明,Character-R1在知识、记忆等方面显著优于现有方法。

原文摘要 · Abstract (English)

Current role-playing agents (RPAs) are typically constructed by imitating surface-level behaviors, but this approach lacks internal cognitive consistency, often causing out-of-character errors in complex situations. To address this, we propose Character-R1, a framework designed to provide comprehensive verifiable reward signals for effective role-aware reasoning, which are missing in recent studies. Specifically, our framework comprises three core designs: (1) Cognitive Focus Reward, which enforces explicit label-based analysis of 10 character elements (e.g., worldview) to structure internal cognition; (2) Reference-Guided Reward, which utilizes overlap-based metrics with reference responses as optimization anchors to enhance exploration and performance; and (3) Character-Conditioned Reward Normalization, which adjusts reward distributions based on character categories to ensure robust optimization across heterogeneous roles. Extensive experiments demonstrate that Character-R1 significantly outperforms existing methods in knowledge, memory and others.

角色扮演强化学习推理机制智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。