黑盒攻击利用视觉扰动破坏多模态记忆,让AI生成错误记忆。
Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents

- 无需访问目标模型,仅通过图像扰动实现攻击
- 攻击成功率超60%,可篡改记忆或注入虚假内容
- 适用于各类记忆架构,暴露系统深层缺陷
多模态AI代理日益依赖持久的长期记忆来锚定生成内容与过往视觉和文本经历。我们发现对视觉数据的无条件信任构成关键漏洞。本文提出Lucid,一种黑盒对抗框架,在严格图像限定威胁模型下攻击多模态记忆管道,无需访问目标大语言模型(MLLM)、目标检索编码器或文本通道。Lucid生成难以察觉的扰动,引发两种不同失效模式:(1) 记忆污染——在上下文内攻击中,对抗图像替换原有图像,其内容因先前文本上下文被强化,可靠地破坏视觉回忆并引导代理走向攻击者设定的叙事;(2) 记忆注入——在无上下文攻击中,对抗图像替换无先前文本支撑的对话回合中的正常图像,导致代理生成受攻击者影响的回应且无记忆纠正信号。我们在多个对话领域及五种黑盒记忆架构上评估了Lucid,包括图结构、LLM摘要式和商用系统。结果显示,污染攻击成功率(ASR)达61.6%,注入攻击达58.4%,揭示多模态记忆管道存在结构性漏洞。
原文摘要 · Abstract (English)
Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual and textual episodes. We show that unconditional trust in visual data creates a critical vulnerability. We propose Lucid, a black-box adversarial framework that compromises multimodal memory pipelines under a strictly image-bounded threat model, requiring no access to the target MLLM, target retrieval encoder, or the text channel. Lucid crafts imperceptible perturbations to enable two distinct failure modes based on the availability of historical context: (1) Memory poisoning, an in-context attack where the adversarial image replaces a benign one whose content is reinforced by prior textual context, reliably corrupting visual recall and steering the agent toward attacker-chosen narratives; (2) Memory injection, an out-of-context attack where the adversarial image replaces a benign one in a conversation turn devoid of prior textual grounding, causing the agent to generate attacker-influenced responses with no corrective signal from memory. We evaluate Lucid across various conversation domains and five black-box memory architectures, including graph-structured, LLM-summarized, and commercially deployed systems. Lucid achieves 61.6% ASR on poisoning and 58.4% ASR on injection, exposing a structural vulnerability in multimodal memory pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。