发现视觉提示注入可让计算机代理被骗,最高欺骗率超50%。
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- 在界面中嵌入恶意指令,通过视觉误导代理执行攻击
- 测试显示现有代理欺骗率最高达51%(CUA)和100%(BUA)
- 适合关注AI安全、多模态代理风险的研究者与开发者
具备完整系统权限的计算机使用代理(CUAs)虽能实现强大任务自动化,但也带来严重的安全与隐私风险,因其可操作文件、访问用户数据并执行任意命令。尽管已有研究聚焦于基于浏览器的代理及HTML层攻击,但对CUAs漏洞的研究仍不充分。本文首次探究视觉提示注入(VPI)攻击,即通过渲染界面中的视觉元素嵌入恶意指令,评估其对CUAs和浏览器使用代理(BUAs)的影响。我们提出VPI-Bench基准,包含跨五个主流平台的306个交互式测试用例,每个用例均在真实环境中部署,并含视觉嵌入的恶意提示。实证研究表明,当前CUAs和BUAs在特定平台上的欺骗率分别高达51%和100%。实验还表明,系统提示防御仅带来有限改善。这些结果凸显了构建鲁棒、上下文感知防御机制的紧迫性,以保障多模态AI代理在现实环境中的安全部署。代码与数据集已开源:https://github.com/cua-framework/agents
原文摘要 · Abstract (English)
Computer-Use Agents (CUAs) with full system access enable powerful task automation but pose significant security and privacy risks due to their ability to manipulate files, access user data, and execute arbitrary commands. While prior work has focused on browser-based agents and HTML-level attacks, the vulnerabilities of CUAs remain underexplored. In this paper, we investigate Visual Prompt Injection (VPI) attacks, where malicious instructions are visually embedded within rendered user interfaces, and examine their impact on both CUAs and Browser-Use Agents (BUAs). We propose VPI-Bench, a benchmark of 306 test cases across five widely used platforms, to evaluate agent robustness under VPI threats. Each test case is a variant of a web platform, designed to be interactive, deployed in a realistic environment, and containing a visually embedded malicious prompt. Our empirical study shows that current CUAs and BUAs can be deceived at rates of up to 51% and 100%, respectively, on certain platforms. The experimental results also indicate that system prompt defenses offer only limited improvements. These findings highlight the need for robust, context-aware defenses to ensure the safe deployment of multimodal AI agents in real-world environments. The code and dataset are available at: https://github.com/cua-framework/agents
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。