首个评估移动端智能体抗视觉干扰攻击的基准,揭示现有模型在动态界面下极易被欺骗。
GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- 在真实安卓模拟器中注入恶意界面元素,动态测试智能体表现
- 当前主流模型在欺骗性界面下失败率超80%,感知与推理均受严重干扰
- 适合安全研究者、移动AI开发者关注,助力构建鲁棒智能体
视觉语言模型(VLMs)正被广泛部署为自主智能体,在移动图形用户界面(GUI)中导航。在包含通知、弹窗和跨应用交互的动态本地环境中,它们面临一种独特且未被充分研究的威胁:环境注入攻击。不同于操纵文本指令的提示攻击,环境注入通过在GUI中插入对抗性界面元素(如欺骗性叠加层或伪造通知),直接破坏智能体的视觉感知,绕过文本防护机制,导致执行偏离,引发隐私泄露、财务损失甚至设备永久损坏。为系统评估此威胁,我们提出GhostEI-Bench,首个在可执行动态环境中评估移动端智能体受环境注入攻击的基准。该基准在全运行状态的安卓模拟器中注入对抗性事件,评估关键风险场景下的性能表现。我们还提出基于判别大模型(judge-LLM)的协议,通过分析智能体动作轨迹与对应截图序列,实现细粒度故障诊断,精准定位感知、识别或推理失败。对前沿智能体的全面实验表明,当前模型对欺骗性环境线索极度脆弱:系统性地无法正确感知和推理被操控的UI。GhostEI-Bench为量化与缓解这一新兴威胁提供了框架,推动更稳健、更安全的具身智能体发展。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly deployed as autonomous agents to navigate mobile graphical user interfaces (GUIs). Operating in dynamic on-device ecosystems, which include notifications, pop-ups, and inter-app interactions, exposes them to a unique and underexplored threat vector: environmental injection. Unlike prompt-based attacks that manipulate textual instructions, environmental injection corrupts an agent's visual perception by inserting adversarial UI elements (for example, deceptive overlays or spoofed notifications) directly into the GUI. This bypasses textual safeguards and can derail execution, causing privacy leakage, financial loss, or irreversible device compromise. To systematically evaluate this threat, we introduce GhostEI-Bench, the first benchmark for assessing mobile agents under environmental injection attacks within dynamic, executable environments. Moving beyond static image-based assessments, GhostEI-Bench injects adversarial events into realistic application workflows inside fully operational Android emulators and evaluates performance across critical risk scenarios. We further propose a judge-LLM protocol that conducts fine-grained failure analysis by reviewing the agent's action trajectory alongside the corresponding screenshot sequence, pinpointing failure in perception, recognition, or reasoning. Comprehensive experiments on state-of-the-art agents reveal pronounced vulnerability to deceptive environmental cues: current models systematically fail to perceive and reason about manipulated UIs. GhostEI-Bench provides a framework for quantifying and mitigating this emerging threat, paving the way toward more robust and secure embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。