让角色扮演模型更像真人:通过意识激活和风格优化提升内在逻辑一致性。
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
- 用角色档案显式引导推理,防止模型忘记自身身份。
- 通过大模型蒸馏使推理风格贴近角色设定,避免过于正式呆板。
- 适合构建有真实感的虚拟角色,如情感陪伴或剧情交互系统。
大型语言模型(LLMs)的发展推动了角色扮演智能体(RPAs)在情感陪伴与虚拟互动等场景的应用。然而,现有RPAs多依赖显式对话数据,缺乏深层的人类式内部思考,导致知识与表达流于表面。尽管大型推理模型(LRMs)可用于模拟角色思维,但直接应用会引发注意力分散(即角色遗忘)与风格漂移(即推理过于正式僵化)。为此,本文提出一种新型角色感知推理(RAR)方法,包含两个关键阶段:角色身份激活(RIA)与推理风格优化(RSO)。RIA在推理过程中显式引入角色档案以缓解注意力分散;RSO则通过LRM蒸馏将推理风格对齐角色与场景特征,有效抑制风格漂移。大量实验证明,所提RAR显著提升了RPAs性能,能有效应对注意力分散与风格漂移问题。
原文摘要 · Abstract (English)
The advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit dialogue data, lacking deep, human-like internal thought processes, resulting in superficial knowledge and style expression. While Large Reasoning Models (LRMs) can be employed to simulate character thought, their direct application is hindered by attention diversion (i.e., RPAs forget their role) and style drift (i.e., overly formal and rigid reasoning rather than character-consistent reasoning). To address these challenges, this paper introduces a novel Role-Aware Reasoning (RAR) method, which consists of two important stages: Role Identity Activation (RIA) and Reasoning Style Optimization (RSO). RIA explicitly guides the model with character profiles during reasoning to counteract attention diversion, and then RSO aligns reasoning style with the character and scene via LRM distillation to mitigate style drift. Extensive experiments demonstrate that the proposed RAR significantly enhances the performance of RPAs by effectively addressing attention diversion and style drift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。