arXiv:2608.21230cs.CRcs.AI2026-08被引 1

虚假记忆可长期存留,现有筛查手段难以防范。

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

  • 用简单生成的假信息污染1.2%数据,准确率从0.85骤降至0.30
  • 写入阶段筛查能发现90%以上诱骗攻击,却无法拦截任何已存假记忆
  • 按来源加权检索无效:强权重会误杀合法但不信任的内容

持久化记忆使虚假信息具有持久性:一旦错误陈述被存储,可在后续匹配会话中反复调用。我们通过单次生成、无指令、无触发器优化的直白假陈述评估此缺陷代价。仅污染LongMemEval语料库的1.2%,准确率即从0.850降至0.300。四阶段写入时筛查管道对间接提示注入的召回率达0.832,但误标1.5%含敏感词的正常文本,却未能拦截360条中毒记忆中的任何一条。这暴露了仅靠内容筛查的局限:区分真假通常需文本之外的外部验证。随后评估溯源加权检索:默认权重与无防御无异(p=0.80),更强权重虽恢复部分性能,但仅通过排除不可信内容实现。在混合溯源语料中,当不可信内容多为良性时,准确率由0.3167升至0.7000;若答案依据本身来自不可信源,证据召回率归零,准确率降至0.0417。在所测相似度条件下,溯源加权项无可用设置:强度足以抵御查询形伪造,也必然压制合法但非可信的证据。因此主张在检索阶段采用有限占用约束,而非加性溯源惩罚,并公开测试工具、语料及汇总报告。

原文摘要 · Abstract (English)

Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it. We measure the cost of this failure mode using plainly worded false assertions generated in a single pass, with no instruction, trigger, or retriever optimization. Poisoning 1.2% of a LongMemEval corpus reduces accuracy from 0.850 to 0.300. A four-stage write-time screening pipeline that reaches 0.832 recall on indirect prompt injection while flagging 1.5% of trigger-word-laden benign text rejects 0 of 360 poisoned memories. We argue this exposes a boundary of content-only screening: distinguishing a false assertion from a true one generally requires external grounding beyond the text itself. We then evaluate provenance-weighted retrieval. The shipped weight is statistically indistinguishable from no defense (p=0.80), while a stronger weight recovers utility only by excluding untrusted content. In a mixed-provenance corpus where untrusted content is mostly benign, accuracy rises from 0.3167 to 0.7000; when the answer-bearing evidence itself arrives untrusted, evidence recall falls to zero and accuracy to 0.0417. Under the measured similarity regime, the additive provenance term has no usable setting: a weight strong enough to resist query-shaped poison is also strong enough to suppress legitimate untrusted evidence. We therefore argue for bounded occupancy constraints at retrieval rather than additive provenance penalties, and release the harnesses, corpora, and aggregate run reports.

记忆安全对抗攻击溯源机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。