评测虚拟角色如何战略性使用记忆,而非仅记住事实。
StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall

- 设计新基准,测试角色在对话中动态调用记忆的能力
- 657个实例中,模型难处理支持性记忆的决策
- 适合研究对话系统与角色智能的开发者
实现虚拟角色的真实对话不仅需要简单记忆和回忆事件,还需战略性地运用记忆以满足事实需求和社交互动。现有相关基准(如记忆增强生成、长对话等)忽略此细节,将记忆视为静态事实库而非可动态部署的资源。为此,我们提出StratMem-Bench,一个评估角色中心对话中战略记忆使用的新型基准。该数据集包含657个实例,虚拟角色需在包含必要、支持性与无关记忆的异构记忆池中作出判断。我们还设计了严格记忆合规度、记忆融合质量、主动丰富度评分和条件无关率等评价指标,用于衡量角色的战略记忆能力。基于最新大语言模型的实验表明:所有模型均能有效区分必要与无关记忆,但在引入支持性记忆后决策能力显著下降。
原文摘要 · Abstract (English)
Achieving realistic human-like conversation for virtual characters requires not only a simple memorization and recall of past events, but also the strategic utilization of memory to meet factual needs and social engagement. Current memory utilization relevant (e.g., memory-augmented generation, long-term dialogue, and etc.) benchmarks overlook this nuance, treating memory primarily as a static repository of facts rather than a dynamic resource to be strategically deployed in dialogues. To address this gap, we design StratMem-Bench, a new benchmark to evaluate strategic memory use in character-centric dialogues. This dataset comprises 657 instances where virtual characters must navigate heterogeneous memory pools containing required, supportive, and irrelevant memories. We also propose a framework with different evaluation metrics including Strict Memory Compliance, Memory Integration Quality, Proactive Enrichment Score and Conditional Irrelevance Rate, to evaluate strategic memory use capabilities of virtual characters. Experiments on StratMem-Bench which leverage the state-of-the-art large language models as virtual characters show that all models perform well at distinguishing between required and irrelevant memories, but struggle once supportive memories are introduced into the decision process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。