arXiv:2605.02199cs.AI2026-05被引 2

提出精确评估长时记忆写作的协议,分离存储与问答干扰。

MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing

论文配图:MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing
图 1 · 摘自论文原文
  • 构建固定包格式的可审计优化问题,明确记忆写入约束条件。
  • 在控制环境下验证记忆质量、有效性与预算选择的独立影响。
  • 提供可复现的生成器、求解器和自然数据包,适合模型评估者使用。

长期大模型代理需在未知未来查询前,将历史交互流压缩为持久记忆。现有评估多依赖最终问答准确率,混淆了记忆写入、检索、提示与阅读推理。本文提出MEMAUDIT,一种针对受限存储预算下长时记忆写入的精确包-预言机评估协议。一个MEMAUDIT包固定经验流、候选记忆表示、存储成本、语义证据单元、未来查询需求及预算,将写入时的记忆选择转化为有限可审计的优化问题,并给出可认证的分母。我们基于凹性-模块化语义覆盖目标,在存储与每经验仅一表示约束下,采用带MILP认证的分支定界法计算精确包最优解。在受控精确包、高有效性压力测试、人工审计的自然支持片段及导出的Mem0、A-Mem、Letta存储中,MEMAUDIT成功分离出表示质量、状态有效性与预算感知选择的影响,而端到端问答无法定位这些效应。成果包括可复用的包生成器、认证求解器、自然包导出、外部系统评分器与缓存的可复现元数据,用于评估记忆写入者在固定存储预算下实际保留的内容。

原文摘要 · Abstract (English)

Long-term LLM agents must compress streams of past interactions into persistent memory before future queries are known. Existing evaluations usually measure final question-answering accuracy, which entangles memory writing with retrieval, prompting, and reader reasoning. We introduce MEMAUDIT, an exact packageoracle evaluation protocol for budgeted long-term memory writing. A MEMAUDIT package fixes an experience stream, candidate memory representations, storage costs, semantic evidence units, future-query requirements, and a budget, turning write-time memory selection into a finite auditable optimization problem with a certified denominator. We instantiate this protocol with a concave-over-modular semantic coverage objective under storage and one-representation-per-experience constraints, and compute exact package optima using branch-and-bound with MILP certification. Across controlled exact packages, validity-heavy stress tests, human-audited natural support slices, and exported Mem0, A-Mem, and Letta stores, MEMAUDIT separates representation quality, validity-state preservation, and budget-aware selection effects that end-to-end QA cannot localize. The resulting artifact provides reusable package generators, certified solvers, natural package exports, external-system scorers, and cached reproducibility metadata for evaluating what memory writers actually preserve under fixed storage budgets.

大模型记忆评估协议存储优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。