arXiv:2502.13053cs.CL2025-02中稿 · ACM MM 2025 Main C…被引 23

AI代理在安卓系统中易被伪装环境干扰,93%成功率攻破决策链。

Evaluating the Robustness of Multimodal Agents Against Active Environmental Injection Attacks

  • 设计新型主动环境注入攻击,利用系统交互漏洞误导多模态代理。
  • 在AndroidWorld基准上实现93%攻击成功率,暴露推理过程脆弱性。
  • 适合安全研究者与移动AI开发者关注,提升智能代理鲁棒性。

随着研究人员持续优化AI代理在操作系统中的任务执行能力,其对环境‘伪造者’的检测能力常被忽视。通过分析代理的操作上下文,我们识别出一项重大威胁:攻击者可将恶意攻击伪装为环境元素,向代理执行过程注入主动干扰以操控其决策。我们定义此新威胁为主动环境注入攻击(AEIA)。聚焦安卓操作系统交互机制,我们评估了AEIA的风险并发现两个关键安全漏洞:(1) 多模态交互界面中的对抗内容注入,攻击者在环境元素中嵌入对抗指令,误导代理决策;(2) 代理任务执行过程中的推理间隙漏洞,使其在推理阶段更易受AEIA攻击。为评估这些漏洞的影响,我们提出AEIA-MN攻击方案,利用移动端操作系统的交互漏洞评估基于多模态大模型(MLLM)代理的鲁棒性。实验结果表明,即使先进MLLM也高度易受攻击,在AndroidWorld基准上结合两种漏洞可实现最高93%的攻击成功率。

原文摘要 · Abstract (English)

As researchers continue to optimize AI agents for more effective task execution within operating systems, they often overlook a critical security concern: the ability of these agents to detect "impostors" within their environment. Through an analysis of the agents' operational context, we identify a significant threat-attackers can disguise malicious attacks as environmental elements, injecting active disturbances into the agents' execution processes to manipulate their decision-making. We define this novel threat as the Active Environment Injection Attack (AEIA). Focusing on the interaction mechanisms of the Android OS, we conduct a risk assessment of AEIA and identify two critical security vulnerabilities: (1) Adversarial content injection in multimodal interaction interfaces, where attackers embed adversarial instructions within environmental elements to mislead agent decision-making; and (2) Reasoning gap vulnerabilities in the agent's task execution process, which increase susceptibility to AEIA attacks during reasoning. To evaluate the impact of these vulnerabilities, we propose AEIA-MN, an attack scheme that exploits interaction vulnerabilities in mobile operating systems to assess the robustness of MLLM-based agents. Experimental results show that even advanced MLLMs are highly vulnerable to this attack, achieving a maximum attack success rate of 93% on the AndroidWorld benchmark by combining two vulnerabilities.

AI安全多模态安卓攻击对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。