让角色扮演模型像人一样思考,更真实地表现角色特质。
Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization

- 用心理学框架拆解角色思考三步:感知互动、共情心理、逻辑构建。
- 新方法在多个数据集上超越现有技术,角色一致性显著提升。
- 适合想打造高拟真角色对话系统的研究者与开发者。
构建能忠实演绎自然语言角色描述的通用角色扮演智能体仍具挑战。当前主流的监督微调方法仅鼓励行为模仿,缺乏深层人类式内在思维,导致泛化能力差。为此,我们提出心理基础的链式思考框架Psy-CoT,将回应前推理分解为三个角色专属步骤:互动感知、心理共情和逻辑建构,使模型从角色描述动态生成思考,而非简单复制表面模式。仅靠结构化推理仍不足;强化学习对角色忠实度至关重要。然而我们发现,在基于大模型的奖励模型下,既能欺骗奖励模型的通用短语与真正角色相关的表达获得相同梯度信号,训练中这种欺骗会累积,误导模型认为两者同样最优。为此,我们提出角色感知策略优化(RAPO),利用角色描述与词元间的互信息,不对称地加权梯度——正优势时放大角色相关词元,负优势时抑制它们。在CoSER、CharacterBench和CharacterEval上的实验表明,Psy-CoT优于现有角色扮演链式思考方法,且RAPO在多种模型规模下持续超越GRPO。
原文摘要 · Abstract (English)
Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant paradigm -- supervised fine-tuning -- encourages behavioral mimicry without deep, human-like internal thought processes, resulting in poor out-of-distribution generalization. Therefore, we propose \textbf{Psy-CoT}, a psychology-grounded chain-of-thought framework that decomposes pre-response reasoning into three role-specific steps -- \emph{Interaction Perception}, \emph{Psychological Empathy}, and \emph{Logical Construction} -- so that the model \emph{thinks dynamically} from the profile rather than merely mimicking surface patterns. While structured reasoning provides a foundation, it alone is insufficient; reinforcement learning is essential to further align the model with character fidelity. However, we observe that under LLM-based reward models, both generic phrases that hack the reward model and genuinely role-specific phrases receive identical gradient signals -- this hacking accumulates over training, misleading the model into treating both as equally optimal choices. To address this, we propose \textbf{Role-Aware Policy Optimization (RAPO)}, which uses profile--token mutual information to weight gradients asymmetrically -- amplifying role-specific tokens under positive advantage while attenuating them under negative advantage. Experiments on CoSER, CharacterBench, and CharacterEval demonstrate that Psy-CoT outperforms existing role-playing CoT methods, and RAPO consistently surpasses GRPO across multiple model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。