用强化学习让智能体自主管理记忆,提升对话长期一致性。
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
- 将记忆管理建模为单智能体强化学习任务,实现端到端控制。
- 在LoCoMo、HaluMem等基准上超越所有现有产品级模型。
- 基于对话数据设计记忆更新奖励机制,适合长时对话场景。
近年来,以人格为中心的记忆系统揭示了多智能体在对话场景中管理人格记忆的强大能力。然而,这些复杂框架常因信息丢失且对不同场景敏感而表现不佳。本文提出DeltaMem,一种基于强化学习的单智能体记忆管理系统,将人格记忆管理视为端到端任务。受人类记忆演化启发,我们构建了一个用户-助手对话数据集,并标注了操作级别的记忆更新标签。在此基础上,引入基于记忆的莱文斯坦距离作为更新奖励,设计专用强化学习框架以增强管理能力。大量实验表明,无论是无需训练还是经强化学习训练的DeltaMem,在多个长期记忆基准(包括LoCoMo、HaluMem、PersonaMem)上均优于所有产品级基线。
原文摘要 · Abstract (English)
Recent advances in persona-centric memory have revealed the powerful capability of multi-agent systems in managing persona memory, especially in conversational scenarios. However, these complex frameworks often suffer from information loss and are fragile across varying scenarios, resulting in suboptimal performance. In this paper, we propose DeltaMem, an agentic memory management system that formulates persona-centric memory management as an end-to-end task within a single-agent setting. To further improve the performance of our agentic memory manager, we draw inspiration from the evolution of human memory and synthesize a user-assistant dialogue dataset along with corresponding operation-level memory updating labels. Building on this, we introduce a novel Memory-based Levenshtein Distance to formalize the memory updating reward, and propose a tailored reinforcement learning framework to further enhance the management capabilities of DeltaMem. Extensive experiments show that both training-free and RL-trained DeltaMem outperform all product-level baselines across diverse long-term memory benchmarks, including LoCoMo, HaluMem, and PersonaMem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。