黑客可借网页内容诱导智能代理执行恶意操作,成功率超80%。
Mind the Web: The Security of Web Use Agents
- 将恶意指令伪装成有用任务提示,利用大模型上下文理解缺陷。
- 攻击成功率达80%以上,跨模型、跨环境具有强迁移性。
- 仅需在公开网站发帖即可发动攻击,连安全机制也难防范。
Web-use agents正被广泛用于自动化复杂浏览器任务,但其强大功能也带来了此前未被探索的攻击面。本文揭示攻击者可通过在网页(如评论、广告)中嵌入恶意内容,诱导代理偏离原任务目标。我们提出任务对齐注入技术,将恶意命令包装为看似有益的任务指引,利用大模型在上下文推理中的根本局限。代理难以维持连贯上下文感知,无法识别隐藏的误导信息。为此,我们构建了三阶段自动化流水线,无需人工标注或昂贵在线交互即可生成有效攻击,即使训练数据有限仍保持高效。该流水线生成的生成器在五种主流代理上测试,针对机密性-完整性-可用性(CIA)三类威胁(包括未经授权摄像头调用、文件外泄、用户冒充、钓鱼、拒绝服务),攻击成功率超过80%,且具备强泛化能力。该攻击在启用内置安全机制的代理上依然有效,仅需在公共网站发布内容即可实施。为此,我们提出包含监督机制、执行约束和任务感知推理在内的综合缓解策略。
原文摘要 · Abstract (English)
Web-use agents are rapidly being deployed to automate complex web tasks with extensive browser capabilities. However, these capabilities create a critical and previously unexplored attack surface. This paper demonstrates how attackers can exploit web-use agents by embedding malicious content in web pages, such as comments, reviews, or advertisements, that agents encounter during legitimate browsing tasks. We introduce the task-aligned injection technique that frames malicious commands as helpful task guidance rather than obvious attacks, exploiting fundamental limitations in LLMs' contextual reasoning. Agents struggle to maintain coherent contextual awareness and fail to detect when seemingly helpful web content contains steering attempts that deviate them from their original task goal. To scale this attack, we developed an automated three-stage pipeline that generates effective injections without manual annotation or costly online agent interactions during training, remaining efficient even with limited training data. This pipeline produces a generator model that we evaluate on five popular agents using payloads organized by the Confidentiality-Integrity-Availability (CIA) security triad, including unauthorized camera activation, file exfiltration, user impersonation, phishing, and denial-of-service. This generator achieves over 80% attack success rate (ASR) with strong transferability across unseen payloads, diverse web environments, and different underlying LLMs. This attack succeed even against agents with built-in safety mechanisms, requiring only the ability to post content on public websites. To address this risk, we propose comprehensive mitigation strategies including oversight mechanisms, execution constraints, and task-aware reasoning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。