用语义进化方法高效生成对抗攻击,突破GUI代理安全瓶颈
EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks
- 在语义层面演化对抗载荷,避开高成本视觉攻击
- 5个目标代理平均攻击成功率85%,1.18~1.71次迭代即成功
- 揭示对权威语义指令的脆弱性,适合安全测试与模型鲁棒性研究
由多模态大语言模型驱动的图形用户界面(GUI)代理日益普及,却易受环境注入攻击(EIAs)威胁。当前红队测试方法面临计算成本高昂和适应性差的问题。我们通过控制实验发现,攻击成功的关键在于语义欺骗而非视觉外观。基于此,提出EVA——一种仅在语义维度演化的对抗框架。EVA采用发现-部署机制,挖掘语言漏洞模式并提炼为通用规则。在五个代表性受害代理上的实验表明,EVA可实现最高85%的攻击成功率,仅需1.18至1.71次迭代即可将良性种子演化为有效攻击。快速收敛揭示了模型潜在表示中密集的语义攻击空间,暴露一个关键对齐悖论:对指令遵循能力的强化训练使代理天然易受权威性、语义欺骗性环境线索影响。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) are increasingly deployed yet vulnerable to Environmental Injection Attacks (EIAs).However, current red-teaming methods are hindered by prohibitive computational costs and limited adaptability. A fundamental question remains unaddressed: does the bottleneck of attack success lie in visual perception or semantic understanding? Through controlled experiments, we observe that semantic deception, rather than visual appearance, serves as the primary determinant of attack success. Based on this insight, we introduce EVA, an evolutionary framework that evolves adversarial payloads exclusively within the semantic dimension. EVA employs a discovery-deployment framework to mine linguistic vulnerability patterns and distill them into generalizable rules. Experimental results across five representative victim agents demonstrate that EVA achieves up to 85\% attack success rate, evolving benign seeds into successful attacks within only 1.18 to 1.71 iterations. This rapid convergence uncovers a dense semantic attack space in the model's latent representation, unveiling a critical alignment paradox: the instruction-following capabilities reinforced by alignment training render agents inherently susceptible to authoritative, semantically deceptive environmental cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。