arXiv:2510.01353cs.AIcs.CL2025-10中稿 · NeurIPS被引 12

评测多平台动态代理环境中的长期记忆与状态追踪能力

MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments

  • 构建跨平台事件交织的动态工作流场景,模拟真实企业环境
  • 主流大模型在长时记忆任务中仅达60%正确率,暴露关键瓶颈
  • 适合研究多智能体协作、记忆增强型代理的学者和工程师

现有上下文与记忆评估多聚焦对话场景,但动态企业环境中记忆能力的评测至关重要。我们提出MEMTRACK,一个用于评估多平台代理环境中长期记忆与状态追踪的基准。该基准通过整合Slack、Linear、Git等多平台的异步事件,构建时间交错、信息噪声大、存在冲突与交叉引用的真实组织工作流。每个实例需处理代码库/文件系统理解与探索,考验记忆获取、选择与冲突解决能力。数据集通过人工专家设计与可扩展的代理生成结合构建,确保生态有效性。引入正确性、效率与冗余度指标,超越传统问答性能。实验显示,当前最先进大模型在长时记忆、跨平台依赖与矛盾解析方面仍面临挑战,最佳GPT-5模型仅达60%正确率。本工作为记忆增强型代理提供可扩展的评估框架,推动复杂组织场景下的多智能体、多平台记忆评测发展。

原文摘要 · Abstract (English)

Recent works on context and memory benchmarking have primarily focused on conversational instances but the need for evaluating memory in dynamic enterprise environments is crucial for its effective application. We introduce MEMTRACK, a benchmark designed to evaluate long-term memory and state tracking in multi-platform agent environments. MEMTRACK models realistic organizational workflows by integrating asynchronous events across multiple communication and productivity platforms such as Slack, Linear and Git. Each benchmark instance provides a chronologically platform-interleaved timeline, with noisy, conflicting, cross-referring information as well as potential codebase/file-system comprehension and exploration. Consequently, our benchmark tests memory capabilities such as acquistion, selection and conflict resolution. We curate the MEMTRACK dataset through both manual expert driven design and scalable agent based synthesis, generating ecologically valid scenarios grounded in real world software development processes. We introduce pertinent metrics for Correctness, Efficiency, and Redundancy that capture the effectiveness of memory mechanisms beyond simple QA performance. Experiments across SoTA LLMs and memory backends reveal challenges in utilizing memory across long horizons, handling cross-platform dependencies, and resolving contradictions. Notably, the best performing GPT-5 model only achieves a 60\% Correctness score on MEMTRACK. This work provides an extensible framework for advancing evaluation research for memory-augmented agents, beyond existing focus on conversational setups, and sets the stage for multi-agent, multi-platform memory benchmarking in complex organizational settings

记忆评测多智能体企业系统长时记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。