受大脑记忆机制启发,提出分两阶段优化记忆的框架,提升长时个性化表现。
Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

- 借鉴前额叶与海马体功能分工,分两阶段优化记忆组织与更新策略
- 在三个基准上显著优于基线,对噪声和规模变化具强鲁棒性
- 适合需要长期记忆的智能体场景,如个性化助手、对话系统
大语言模型智能体需长期用户记忆以实现持续个性化,但有限的上下文窗口阻碍了对长期交互中演变偏好的追踪。现有记忆系统主要依赖静态的手工更新规则;尽管基于强化学习的智能体可学习记忆更新,但稀疏的结果奖励提供弱监督,导致长时优化不稳定。受记忆图式理论及前额叶与海马体功能分工启发,我们提出 MemCoE——一种认知启发的两阶段优化框架,用于学习记忆应如何组织以及更新哪些信息。第一阶段通过对比反馈(视为文本梯度)优化全局记忆指导原则;第二阶段利用该指导原则定义结构化过程奖励,进行多轮强化学习,以学习遵循指导原则的记忆演化策略。我们在三个个性化记忆基准上评估,涵盖显式/隐式偏好及不同规模与噪声水平,结果一致优于强基线,具备良好鲁棒性、可迁移性和效率。
原文摘要 · Abstract (English)
Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving preferences over long interactions. Existing memory systems mainly rely on static, hand-crafted update rules; although reinforcement learning (RL)-based agents learn memory updates, sparse outcome rewards provide weak supervision, resulting in unstable long-horizon optimization. Drawing on memory schema theory and the functional division between prefrontal regions and hippocampus regions, we introduce MemCoE, a cognition-inspired two-stage optimization framework that learns how memory should be organized and what information to update. In the first stage, we propose Memory Guideline Induction to optimize a global guideline via contrastive feedback interpreted as textual gradients; in the second stage, Guideline-Aligned Memory Policy Optimization uses the induced guideline to define structured process rewards and performs multi-turn RL to learn a guideline-following memory evolution policy. We evaluate on three personalization memory benchmarks, covering explicit/implicit preference and different sizes and noise, and observe consistent improvements over strong baselines with favorable robustness, transferability, and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。