通过操控视觉定位实现隐蔽后门攻击,威胁智能界面代理安全
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- 在预训练阶段注入毒化数据,误导模型将文本指令误映射到错误界面元素
- 仅需5%毒化数据即可成功劫持视觉定位,且触发信号对人眼不可见
- 攻击可跨平台泛化,即使经过清理微调仍有效,适合研究安全漏洞者
基于大视觉语言模型的图形用户界面(GUI)代理正推动人机交互自动化,能自主操作手机等设备完成复杂任务。然而其与个人设备深度集成,带来显著安全隐患,其中后门攻击尚少被探讨。本文揭示,GUI代理将文本计划映射至界面元素的视觉定位环节存在漏洞,可被用于新型后门攻击。提出VisualTrap方法,通过在视觉定位预训练阶段注入毒化数据,误导代理将文本指令关联至触发位置而非目标元素。实验证明,该攻击仅需5%毒化数据即可有效劫持视觉定位,触发信号对人眼不可见;且攻击可泛化至下游任务,即便经过清洁微调仍保持效果。此外,训练于移动端/网页端的触发器可跨平台迁移至桌面环境。这些发现凸显了对GUI代理后门风险亟需深入研究。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) agents powered by Large Vision-Language Models (LVLMs) have emerged as a revolutionary approach to automating human-machine interactions, capable of autonomously operating personal devices (e.g., mobile phones) or applications within the device to perform complex real-world tasks in a human-like manner. However, their close integration with personal devices raises significant security concerns, with many threats, including backdoor attacks, remaining largely unexplored. This work reveals that the visual grounding of GUI agent-mapping textual plans to GUI elements-can introduce vulnerabilities, enabling new types of backdoor attacks. With backdoor attack targeting visual grounding, the agent's behavior can be compromised even when given correct task-solving plans. To validate this vulnerability, we propose VisualTrap, a method that can hijack the grounding by misleading the agent to locate textual plans to trigger locations instead of the intended targets. VisualTrap uses the common method of injecting poisoned data for attacks, and does so during the pre-training of visual grounding to ensure practical feasibility of attacking. Empirical results show that VisualTrap can effectively hijack visual grounding with as little as 5% poisoned data and highly stealthy visual triggers (invisible to the human eye); and the attack can be generalized to downstream tasks, even after clean fine-tuning. Moreover, the injected trigger can remain effective across different GUI environments, e.g., being trained on mobile/web and generalizing to desktop environments. These findings underscore the urgent need for further research on backdoor attack risks in GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。