arXiv:2606.29788cs.LG2026-06被引 1

多模态智能体删了文字却留了图像漏洞,信息仍可被恢复。

MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory

论文配图:MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory
图 1 · 摘自论文原文
  • 构建信息溯源图,分析记忆删除的失效路径
  • 图像泄露率达12.0%,47%漏出无法通过文本还原
  • 内容感知删除能将残留降至2.0%,适合隐私保护研究者

当多模态AI智能体被要求遗忘某条信息时,现有记忆系统通常仅删除文本条目并报告成功。我们发现,该信息仍可通过保留的用户图像被恢复,包括与不同事实相关联的图像,因为视觉语言模型(VLM)在推理时会利用隐含视觉线索。为此,我们提出信息溯源图(IPG),对记忆表示按删除可行性进行分类,揭示删除失败存在多重通道。我们的基准测试MemLeak衡量了删除链中的泄露情况:直接探测删除能力系统仅泄露<1%,保留的关联文本使恢复率达18.3%,而保留图像导致12.0%恢复率(盲基线0.0%,假阳性率0.3%),其中47%的图像泄露无法通过文本还原。内容感知语义删除可将图像残留降至2.0%。该残余现象出现在多个VLM、生产级记忆系统及真实Unsplash授权照片中。双标注员人工验证(kappa=0.88)确认了判断可靠性。

原文摘要 · Abstract (English)

When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success. We find that the fact can remain recoverable from retained user images, including images tagged to entirely different facts, because VLMs use implicit visual cues at inference time. We introduce the Information Provenance Graph (IPG), a taxonomy that classifies memory representations by deletion affordance. The IPG reveals that deletion fails through multiple channels. Our benchmark, MemLeak, measures this across a deletion cascade: direct probing of deletion-capable systems yields <1%, but retained correlated text enables 18.3% recovery, and retained images enable 12.0% recovery (0.0% blind baseline, 0.3% FPR) -- with 47% of image leaks not text-recoverable. Content-aware semantic deletion reduces the image residual to 2.0%. The residual appears across multiple VLMs, a production memory system, and real Unsplash-licensed photographs. Dual-annotator human validation (kappa = 0.88) confirms judge reliability.

多模态隐私泄露记忆系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。