arXiv:2608.12476cs.AI2026-08

为长时智能体设计可审计的记忆系统,确保信息更新不回溯、删除不可恢复。

Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

  • 引入源绑定记忆模型,记录信息来源与生命周期状态
  • 在3600个测试案例中完全匹配正确结果,99.8%以上故障修复率
  • 适合需要严格数据追溯与安全释放的系统级应用

长期智能体记忆传统上是选择-存储-检索模式,但检索无法判断矛盾、覆盖、撤回、删除或过期记录是否仍可支持当前主张。本文提出受控持久记忆(GPM),一种可审计的双时态状态转换模型,具备源绑定准入、衍生生命周期状态、当前公开屏障及失败闭合结构化释放机制。五个可执行条款涵盖账本完整性、源绑定、冲突隔离、撤回或删除后不可恢复性,以及在单一验证头下的精确主张关闭。在预设哈希冻结的3,600例GPM-ReleaseBench测试中,GPM完全匹配所有完整结果;三种简单策略中最优者在1,800/3,600案例中成功,且在50%违规案例中实现未匹配释放。独立密封端到端服务评估覆盖八个查询类别,公开披露的V3版本中,受控路径在2,400/2,400集群上正确,而未受控本地Qwen2.5-7B仅正确600/2,400;其修复全部1,800个基线失败,无回归(单边95%置信下界分别为99.875%和99.834%)。后续V5重新封存版本在中英文指令分支均实现2,400/2,400正确率,含生成日期锁定且冻结后不可修改。生产无关有限模型探索331,776种语义状态与1,990,656种查询状态,未发现全合约反例;10万条三引擎差分轨迹亦零错配。这些为合约边界与实现结果,非开放世界模型精度或真实性的证据。密封服务中的受控输出为确定性服务结果;7B对比结果仅为未受控参照,不代表语言模型本身达到完美准确。

原文摘要 · Abstract (English)

Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitemporal state-transition model with source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release. Five executable clauses cover ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and exact claim closure over a fresh view at one verified head. On a prespecified hash-frozen 3,600-case GPM-ReleaseBench, GPM matches all complete outcomes; the strongest of three intentionally simple complete policies matches 1,800/3,600 and makes unmatched releases on 50% of violation cases. A separate sealed end-to-end service evaluation exercises real ingestion and release across eight query families. In its publicly disclosed V3 arm, the governed lane is correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B; it repairs all 1,800 baseline failures with no regression (one-sided 95% lower bounds 99.875% and 99.834%). A later V5 reseal over Chinese- and English-command arms, with generation-date pinning and no post-freeze reducer amendment, again obtains 2,400/2,400 per arm. A production-code-independent finite model explores 331,776 semantic and 1,990,656 query states without a full-contract counterexample, and a 100,000-trace three-engine differential yields zero mismatches. These are bounded contract and implementation results, not open-world model accuracy or evidence of world truth. Governed answers in the sealed service evaluation are deterministic service outputs; the 7B result is the ungoverned comparison, not a claim that a language model itself became perfectly accurate.

智能体记忆可审计系统状态管理安全释放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。