剖析智能体记忆系统缺陷,揭示评估与性能短板
Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations
- 按四种记忆结构建立智能体记忆分类体系
- 发现评估指标失真、模型依赖性强、延迟开销大等关键问题
- 适合关注LLM长时推理与系统优化的研究者
智能体记忆系统使大语言模型代理能在长时间交互中保持状态,支持超出固定上下文窗口的长周期推理与个性化。尽管架构发展迅速,现有实证基础仍较薄弱:基准测试规模不足,评估指标与语义效用不匹配,不同主干模型表现差异显著,系统级开销常被忽略。本文从架构与系统双重视角对智能体记忆进行结构化分析,首先基于四种记忆结构提出简洁分类体系;接着分析当前系统的四大痛点:基准饱和效应、指标有效性与评判敏感性、主干模型依赖性,以及记忆维护带来的延迟与吞吐量开销。通过关联记忆结构与实证限制,阐明现有系统为何常未能兑现理论潜力,并提出更可靠评估与可扩展系统设计的方向。
原文摘要 · Abstract (English)
Agentic memory systems enable large language model (LLM) agents to maintain state across long interactions, supporting long-horizon reasoning and personalization beyond fixed context windows. Despite rapid architectural development, the empirical foundations of these systems remain fragile: existing benchmarks are often underscaled, evaluation metrics are misaligned with semantic utility, performance varies significantly across backbone models, and system-level costs are frequently overlooked. This survey presents a structured analysis of agentic memory from both architectural and system perspectives. We first introduce a concise taxonomy of MAG systems based on four memory structures. Then, we analyze key pain points limiting current systems, including benchmark saturation effects, metric validity and judge sensitivity, backbone-dependent accuracy, and the latency and throughput overhead introduced by memory maintenance. By connecting the memory structure to empirical limitations, this survey clarifies why current agentic memory systems often underperform their theoretical promise and outlines directions for more reliable evaluation and scalable system design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。