让聊天机器人像人一样有结构化记忆,提升长期对话能力。
SaliMory: Orchestrating Cognitive Memory for Conversational Agents

- 用分阶段奖励和对比优化,让模型自主管理记忆的筛选、整合与调用。
- 记忆错误减少三分之一,端到端准确率提升超10%,个性化表现翻倍。
- 适合做长期陪伴型对话系统的研究者或开发者参考。
作为终身伴侣的对话代理需在所有交互中保持持久记忆。然而,单纯扩大上下文窗口并结合原始检索会降低推理质量;而通过标准强化学习训练记忆代理,则在多阶段流程中引发严重的信用分配瓶颈。为此,我们提出SALIMORY框架,让单一语言模型管理用户事实、偏好与工作记忆等认知结构化记忆。通过引入分阶段过程奖励与奖励分解对比优化,SALIMORY为不同记忆操作(选择性过滤、整合、线索驱动召回)提供独立监督,实现端到端训练。SALIMORY将记忆相关错误减少三分之一,在端到端准确率上超越当前最优方法超过10%,个性化表现提升逾一倍。
原文摘要 · Abstract (English)
Conversational agents that serve as lifelong companions must maintain persistent memory across all interactions. However, simply expanding context windows with raw retrieval degrades reasoning quality, while training memory agents via standard reinforcement learning creates a severe credit assignment bottleneck in a multi-stage pipeline. To solve this, we introduce SALIMORY, a framework that trains a single language model to manage a cognitively-structured memory-spanning user facts, preferences, and working memory. By introducing a hierarchical stage-wise process reward and reward-decomposed contrastive refinement, SALIMORY provides isolated supervision for distinct memory operations (selective filtering, consolidation, and cue-driven recall) end-to-end. SALIMORY cuts memory-attributed failures by one-third, outperforms the state-of-the-art by over 10% in end-to-end accuracy, and more than doubles the Good Personalization rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。