用强化学习生成可控制的恶意指令,攻击网页智能体
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
- 用强化学习训练对抗性提示器,根据目标智能体反馈优化攻击指令
- 在多种网页任务中对GPT-4驱动的智能体攻击成功率高,且隐蔽性强
- 揭示现有防御手段效果有限,适合安全研究人员和模型开发者
基于基础模型的智能体正被广泛用于自动化复杂任务,提升效率与生产力。然而,其对敏感资源的访问权限及自主决策能力也带来了显著安全风险,一旦遭受攻击可能造成严重后果。为系统性发现这些漏洞,我们提出AdvAgent——一种针对网页智能体的黑盒红队攻击框架。不同于现有方法,AdvAgent采用基于强化学习的流程,训练一个对抗性提示器模型,通过黑盒智能体的反馈来优化对抗性提示。精心设计的攻击指令能有效利用智能体弱点,同时保持隐蔽性和可控性。大量实验表明,AdvAgent在多种网页任务中对当前最先进的GPT-4基智能体均表现出高攻击成功率。此外,我们发现现有基于提示的防御措施仅提供有限保护,使智能体仍易受本框架攻击。这些结果揭示了当前网页智能体的关键漏洞,凸显加强防御机制的紧迫性。代码已开源:https://ai-secure.github.io/AdvAgent/
原文摘要 · Abstract (English)
Foundation model-based agents are increasingly used to automate complex tasks, enhancing efficiency and productivity. However, their access to sensitive resources and autonomous decision-making also introduce significant security risks, where successful attacks could lead to severe consequences. To systematically uncover these vulnerabilities, we propose AdvAgent, a black-box red-teaming framework for attacking web agents. Unlike existing approaches, AdvAgent employs a reinforcement learning-based pipeline to train an adversarial prompter model that optimizes adversarial prompts using feedback from the black-box agent. With careful attack design, these prompts effectively exploit agent weaknesses while maintaining stealthiness and controllability. Extensive evaluations demonstrate that AdvAgent achieves high success rates against state-of-the-art GPT-4-based web agents across diverse web tasks. Furthermore, we find that existing prompt-based defenses provide only limited protection, leaving agents vulnerable to our framework. These findings highlight critical vulnerabilities in current web agents and emphasize the urgent need for stronger defense mechanisms. We release code at https://ai-secure.github.io/AdvAgent/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。