arXiv:2604.02623cs.CRcs.AI2026-04被引 14

攻击者仅通过观察网页就能长期污染智能代理的记忆。

Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents

  • 利用环境观察实现跨会话、跨网站记忆污染
  • 最高攻击成功率32.5%,压力环境下提升8倍
  • 适合关注AI浏览器安全的研究者与开发者

基于大模型的网络代理依赖记忆来个性化任务,但也因此形成持久攻击面。现有研究多假设攻击者可直接注入或共享内存,但本文提出更现实的威胁模型:仅通过环境观察即可污染记忆。我们提出环境注入轨迹式代理记忆投毒(eTAMP),首次实现无需直接内存访问的跨会话、跨网站攻击。一次恶意页面浏览即可无声污染代理记忆,并在后续不同网站任务中激活,绕过权限防御。在(视觉)WebArena上的实验表明:eTAMP对GPT-5-mini、GPT-5.2、GPT-OSS-120B的攻击成功率分别达32.5%、23.4%、19.5%。此外发现“挫败感利用”现象:当代理遭遇点击失效或文本乱码等环境压力时,攻击成功率最高提升8倍。值得注意的是,能力更强的模型未必更安全,如GPT-5.2虽表现优异却仍具显著漏洞。随着OpenClaw、ChatGPT Atlas、Perplexity Comet等AI浏览器兴起,亟需针对环境注入型记忆投毒建立防御机制。

原文摘要 · Abstract (English)

Memory makes LLM-based web agents personalized, powerful, yet exploitable. By storing past interactions to personalize future tasks, agents inadvertently create a persistent attack surface that spans websites and sessions. While existing security research on memory assumes attackers can directly inject into memory storage or exploit shared memory across users, we present a more realistic threat model: contamination through environmental observation alone. We introduce Environment-injected Trajectory-based Agent Memory Poisoning (eTAMP), the first attack to achieve cross-session, cross-site compromise without requiring direct memory access. A single contaminated observation (e.g., viewing a manipulated product page) silently poisons an agent's memory and activates during future tasks on different websites, bypassing permission-based defenses. Our experiments on (Visual)WebArena reveal two key findings. First, eTAMP achieves substantial attack success rates: up to 32.5% on GPT-5-mini, 23.4% on GPT-5.2, and 19.5% on GPT-OSS-120B. Second, we discover Frustration Exploitation: agents under environmental stress become dramatically more susceptible, with ASR increasing up to 8 times when agents struggle with dropped clicks or garbled text. Notably, more capable models are not more secure. GPT-5.2 shows substantial vulnerability despite superior task performance. With the rise of AI browsers like OpenClaw, ChatGPT Atlas, and Perplexity Comet, our findings underscore the urgent need for defenses against environment-injected memory poisoning.

AI安全记忆投毒智能代理环境攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。