MUZZLE自动发现网页代理的隐蔽攻击面并动态发起针对性攻击。
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- 基于代理行为轨迹自动定位高风险注入点
- 在4个应用中发现44种新攻击,覆盖10类恶意目标
- 适合安全研究人员和开发人员评估代理鲁棒性
基于大语言模型的网页代理正被广泛用于自动化复杂在线任务,但其设计使其易受嵌入不可信网页内容的间接提示注入攻击,攻击者可劫持代理行为并违背用户意图。现有评估方法依赖固定攻击模板、人工选择注入点或局限场景,难以模拟真实环境中自适应攻击。本文提出MUZZLE,一个自动化代理式评估框架,通过分析代理执行轨迹自动识别高显著性注入表面,并动态生成上下文感知的恶意指令,针对机密性、完整性与可用性等目标发起攻击。不同于以往方法,MUZZLE根据代理实际执行路径调整攻击策略,并利用失败反馈迭代优化攻击。我们在多种网页应用、用户任务和代理配置下评估该框架,结果表明其能以极低人工干预实现对网页代理安全性的自动适应性评估。共发现44种新攻击,涉及4个应用和10类恶意目标,涵盖不同LLM及代理架构;还识别出3种跨应用提示注入攻击及一种定制化钓鱼场景。
原文摘要 · Abstract (English)
Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf. While these agents offer powerful capabilities, their design exposes them to indirect prompt injection attacks embedded in untrusted web content, enabling adversaries to hijack agent behavior and violate user intent. Despite growing awareness of this threat, existing evaluations rely on fixed attack templates, manually selected injection surfaces, or narrowly scoped scenarios, limiting their ability to capture realistic, adaptive attacks encountered in practice. We present MUZZLE, an automated agentic framework for evaluating the security of web agents against indirect prompt injection attacks. MUZZLE utilizes the agent's trajectories to automatically identify high-salience injection surfaces, and adaptively generate context-aware malicious instructions that target violations of confidentiality, integrity, and availability. Unlike prior approaches, MUZZLE adapts its attack strategy based on the agent's observed execution trajectory and iteratively refines attacks using feedback from failed executions. We evaluate MUZZLE across diverse web applications, user tasks, and agent configurations, demonstrating its ability to automatically and adaptively assess the security of web agents with minimal human intervention. Our results show that MUZZLE effectively discovers 44 new attacks on 4 web applications with 10 adversarial objectives that violate confidentiality, availability, or privacy properties across different LLMs and agent scaffolds. MUZZLE also identifies novel attack strategies, including 3 cross-application prompt injection attacks and an agent-tailored phishing scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。