让聊天角色更像自己:解决大模型角色扮演中的风格坍缩问题
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

- 分离任务奖励与风格奖励,避免优化冲突
- 根据角色复杂度动态调整训练约束,提升个性化表现
- 用通用回复作负样本,防止模型趋同于平庸输出
近期强化学习进展,特别是群体相对策略优化(GRPO),显著提升了大语言模型的推理能力。然而,将这类以任务为中心的优化方法应用于角色扮演智能体时,常导致角色特征失真和风格坍缩,因其优先考虑上下文实用性而非角色一致性。为此,我们提出面向角色感知推理的角色中心型群体相对策略优化(CRPO)框架,通过三种机制重对齐强化学习目标:解耦任务逻辑与风格奖励以缓解梯度冲突;根据角色复杂度动态调整优化约束;利用通用回复作为负向基线,防止模型退化至共通分布。大量实验证明,CRPO在一致性、情感表达等方面优于现有方法。
原文摘要 · Abstract (English)
Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning capabilities of Large Language Models. However, applying these problem-centric optimization methods to role-playing agents often leads to a loss of character fidelity and style collapse, as they prioritize context-specific utility over persona alignment. To address this, we propose Character-Centric Group Relative Policy Optimization (CRPO), a framework designed to realign RL objectives with the role-playing task. CRPO improves character distinctiveness through three mechanisms: decoupling task logic from stylistic rewards to resolve gradient conflicts, dynamically adapting optimization constraints based on character complexity, and utilizing generic responses as negative baselines to prevent the model from reverting to a common distribution. Extensive experiments demonstrate that CRPO outperforms existing methods in consistency, emotion and others.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。