arXiv:2507.17147cs.CL2025-07EMNLP被引 6

让大模型像人一样思考后再回应,提升角色扮演一致性。

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards

  • 采用认知-回应双阶段推理,融合外部情境与内部自我意识。
  • 在多个基准上优于现有方法,角色一致性显著提升。
  • 适合需要稳定角色表现的对话系统与虚拟助手场景。

角色扮演语言代理(RPLAs)已成为大语言模型的重要应用方向。现有方法多依赖提示工程或监督微调来模拟特定场景下的角色行为,但常忽略驱动行为背后的认知机制。受认知心理学启发,我们提出新型RPLA CogDual,采用‘先认知后回应’的推理范式。通过联合建模外部情境感知与内部自我意识,CogDual生成的回答在角色一致性和上下文对齐性上均有提升。为进一步优化性能,我们设计了两种通用奖励机制,结合强化学习用于开放域文本生成。在CoSER、Cross-MR和LifeChoice等多个基准上的实验表明,CogDual持续优于现有基线,且在多种角色扮演任务中表现出良好泛化能力。

原文摘要 · Abstract (English)

Role-Playing Language Agents (RPLAs) have emerged as a significant application direction for Large Language Models (LLMs). Existing approaches typically rely on prompt engineering or supervised fine-tuning to enable models to imitate character behaviors in specific scenarios, but often neglect the underlying \emph{cognitive} mechanisms driving these behaviors. Inspired by cognitive psychology, we introduce \textbf{CogDual}, a novel RPLA adopting a \textit{cognize-then-respond } reasoning paradigm. By jointly modeling external situational awareness and internal self-awareness, CogDual generates responses with improved character consistency and contextual alignment. To further optimize the performance, we employ reinforcement learning with two general-purpose reward schemes designed for open-domain text generation. Extensive experiments on the CoSER benchmark, as well as Cross-MR and LifeChoice, demonstrate that CogDual consistently outperforms existing baselines and generalizes effectively across diverse role-playing tasks.

角色扮演认知建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。