arXiv:2603.19313cs.CLcs.AI2026-03ACL被引 2

让大模型像演员一样靠记忆扮演角色,提升长期对话一致性。

Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs

论文配图:Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
图 1 · 摘自论文原文
  • 将角色知识视为模型内部记忆,仅凭对话上下文自动调用。
  • 小模型用新方法可达到大模型表现,最大提升37%效果。
  • 提供中英文双语评测基准,适合角色扮演和对话系统研究者。

大模型在长时间开放式对话中难以保持角色一致性,常因无法回忆或准确应用角色知识而失真。为此,本文提出“记忆驱动型角色扮演”范式,借鉴斯坦尼斯拉夫斯基的“情感记忆”理论,将角色知识视为模型内部记忆库,要求仅依据对话上下文完成记忆检索与应用,从而严格检验知识的深度与自主使用能力。基于该范式,贡献包括:(1) MREval——细粒度评估框架,涵盖锚定、回忆、边界控制与角色演绎四类能力;(2) MRPrompt——引导结构化记忆检索与生成的提示架构;(3) MRBench——中英双语基准,支持细粒度诊断。该范式对12个大模型进行四阶段角色扮演能力评估。实验表明,采用MRPrompt后,小模型(如Qwen3-8B)性能可媲美大型闭源模型(如Qwen3-Max和GLM-4.7),且上游记忆能力提升直接带来下游响应质量改善,验证了分阶段理论的有效性。

原文摘要 · Abstract (English)

A core challenge for faithful LLM role-playing is sustaining consistent characterization throughout long, open-ended dialogues, as models frequently fail to recall and accurately apply their designated persona knowledge without explicit cues. To tackle this, we propose the Memory-Driven Role-Playing paradigm. Inspired by Stanislavski's "emotional memory" acting theory, this paradigm frames persona knowledge as the LLM's internal memory store, requiring retrieval and application based solely on dialogue context, thereby providing a rigorous test of depth and autonomous use of knowledge. Centered on this paradigm, we contribute: (1) MREval, a fine-grained evaluation framework assessing four memory-driven abilities - Anchoring, Recalling, Bounding, and Enacting; (2) MRPrompt, a prompting architecture that guides structured memory retrieval and response generation; and (3) MRBench, a bilingual (Chinese/English) benchmark for fine-grained diagnosis. The novel paradigm provides a comprehensive diagnostic for four-staged role-playing abilities across 12 LLMs. Crucially, experiments show that MRPrompt allows small models (e.g., Qwen3-8B) to match the performance of much larger closed-source LLMs (e.g., Qwen3-Max and GLM-4.7), and confirms that upstream memory gains directly enhance downstream response quality, validating the staged theoretical foundation.

角色扮演记忆机制大模型评估提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。