arXiv:2603.01966cs.CLcs.AI2026-03中稿 · ICLR被引 16

构建交互式记忆评估框架,提升长对话中助手的记忆能力测试与优化

AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations

  • 通过结构化数据采样生成用户画像与状态演化轨迹,实现动态评估
  • 实验发现现有记忆系统存在显著性能差距,如RAG与长上下文模型表现不佳
  • 适合研究对话记忆、个性化助理及自进化系统的研究者使用

基于大模型的助手在长时对话中需要有效的记忆管理,但当前方法在训练与评估方面面临挑战。现有记忆基准依赖静态、离线的数据作为上下文,限制了评估的可靠性与可扩展性。为此,我们提出AMemGym,一个交互式环境,支持记忆驱动个性化系统的在线策略评估与优化。AMemGym采用结构化数据采样,预定义用户画像、状态相关问题与状态演化路径,可低成本生成高质量、评估对齐的交互数据。由大模型模拟用户通过角色扮演揭示隐含状态,同时保持结构化状态一致性。基于结构化数据的综合指标,既可用于评估也可用于优化助手表现。大量实验揭示现有记忆系统(如RAG、长上下文大模型、代理型记忆)存在显著性能差距及其成因。AMemGym不仅能有效筛选竞争性方案,还可推动记忆管理策略的自我演进。通过将结构化状态演化与自由形式交互结合,本框架为对话智能体的记忆能力发展提供了可扩展、诊断性强的环境。

原文摘要 · Abstract (English)

Long-horizon interactions between users and LLM-based assistants necessitate effective memory management, yet current approaches face challenges in training and evaluation of memory. Existing memory benchmarks rely on static, off-policy data as context, limiting evaluation reliability and scalability. To address these gaps, we introduce AMemGym, an interactive environment enabling on-policy evaluation and optimization for memory-driven personalization. AMemGym employs structured data sampling to predefine user profiles, state-dependent questions, and state evolution trajectories, enabling cost-effective generation of high-quality, evaluation-aligned interactions. LLM-simulated users expose latent states through role-play while maintaining structured state consistency. Comprehensive metrics based on structured data guide both assessment and optimization of assistants. Extensive experiments reveal performance gaps in existing memory systems (e.g., RAG, long-context LLMs, and agentic memory) and corresponding reasons. AMemGym not only enables effective selection among competing approaches but also can potentially drive the self-evolution of memory management strategies. By bridging structured state evolution with free-form interactions, our framework provides a scalable, diagnostically rich environment for advancing memory capabilities in conversational agents.

对话记忆评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。