编码智能体的记忆管理需考虑语义差异,否则评估会失真。
Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents
- 按记忆内容的语义类型区分管理策略
- 不同语义对象留存与压缩行为差异显著
- 评估需超越令牌预算,关注实际上下文交付
智能体工作记忆具有语义异质性,如指令、产物、工具输出和自生成状态等对象在语义角色、大小、保留性和表征上各不相同。本研究基于55条归档编码智能体轨迹,发现语义不同的记忆对象表现出显著不同的保留与压缩行为,这促使采用语义感知的记忆管理机制。我们考察了两种语义感知策略:对象感知压缩策略与基于检索的策略。评估表明,校准收益未必可迁移至未见任务,且相同令牌预算不等于等效上下文交付或管理开销。真实系统回放进一步揭示了仅凭名义预算无法捕捉的服务瓶颈。这些结果表明,语义结构对智能体记忆至关重要,而评估记忆管理策略不能仅依赖名义令牌预算。我们据此归纳出四个层次:存储状态、交付上下文、管理开销与任务或过程结果。
原文摘要 · Abstract (English)
Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated state play different semantic roles and exhibit different size, retention, and representation profiles. Recent work has begun to explore memory-management mechanisms that account for such heterogeneity. This work focuses on semantic heterogeneity and studies how it should shape the management and evaluation of working memory in coding agents. Across 55 archived coding-agent trajectories, we find that semantically different working-memory objects exhibit distinct retention and compression behavior. This heterogeneity motivates semantically informed memory management. We study two semantically informed strategies: an object-aware compression policy and a retrieval-based policy. Their evaluation shows that calibration gains may not transfer to held-out tasks, and that equal token budgets do not imply equal delivered context or management cost. A real-system replay further exposes serving limits that nominal budgets alone do not capture. Together, these results show why semantic structure matters for agent working memory and why evaluating memory-management strategies requires more than a nominal token budget. We organize these lessons into four levels: stored state, delivered context, management work, and task or process outcome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。