arXiv:2606.10742cs.CRcs.LG2026-06

攻击网页智能体记忆系统,用图文组合骗其长期执行恶意任务。

MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents

论文配图:MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents
图 1 · 摘自论文原文
  • 用图文协同伪造记忆,让恶意内容高概率被调用。
  • 攻击成功率最高达99.15%,且不影响正常功能。
  • 无需改模型,可跨架构复用,适合安全研究者关注。

外部记忆已成为现代网页智能体的核心组件,通过检索过往经验实现长时序推理。然而,这一范式引入了关键漏洞:恶意内容注入记忆后可被持续召回并反复影响智能体行为。本文首次系统研究多模态记忆投毒这一被忽视但现实存在的攻击面。提出MemVenom——一种统一的黑盒攻击框架,通过协调文本与图像证据,污染图结构的外部记忆。方法包含两阶段:(1) 触发条件引导的检索攻击,确保恶意记忆高概率被调用;(2) 检索后攻击诱导,利用对抗扰动与隐蔽OCR注入覆盖原始用户目标。相比以往仅作用于提示或纯文本记忆的攻击,本方法实现持久、可复用、目标无关的攻击,无需修改模型参数或重新优化恶意任务。在多个网页智能体框架和视觉-语言模型上实验表明,MemVenom达成强端到端攻击成功率,对良性性能影响极小,最高达99.15%(针对GPT-5家族),且在不同架构与模型规模间具有良好迁移性。

原文摘要 · Abstract (English)

External memory has become a core component of modern web agents, enabling long-horizon reasoning through the retrieval of past experiences. However, this paradigm introduces a critical vulnerability: malicious content injected into memory can be persistently recalled and repeatedly influence agent behavior. In this work, we identify and systematically study multimodal memory poisoning, an overlooked yet practical attack surface in web-agent systems. We propose MemVenom, a unified black-box attack framework that poisons graph-structured external memory with coordinated text-image evidence. Our method consists of a two-stage design: (1) a trigger-conditioned retrieval attack that ensures high-probability recall of malicious memory, and (2) a post-retrieval attack induction that leverages adversarial perturbations and stealthy OCR injection to override the original user objective. Unlike prior attacks that operate on prompts or text-only memory, our approach enables persistent, reusable, and goal-agnostic attacks without modifying model parameters or re-optimizing malicious tasks. Experiments across multiple web-agent frameworks and vision-language models demonstrate that MemVenom achieves strong end-to-end attack success with minimal impact on benign performance, reaching up to 99.15% on GPT-5-family web agents, while transferring effectively across architectures and model scales.

智能体安全记忆投毒多模态攻击黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。