用人类看不见的视觉陷阱,劫持手机智能体的指令执行
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
- 仅在智能体交互时暴露恶意内容,人眼无法察觉
- 单次优化生成攻击提示,绕过安全过滤器成功率82.5%
- 适合研究移动智能体安全或对抗攻击的学者
大型视觉语言模型(LVLMs)赋予移动智能体自主能力,但其在真实移动部署环境下的安全性仍缺乏深入研究。尽管智能体易受视觉提示注入攻击,但无需系统权限的隐蔽攻击仍具挑战性,因现有方法依赖持续可见的视觉篡改。我们发现人类与智能体交互存在显著差异:自动化智能体产生的触控信号几乎为零。基于此,提出仅针对智能体的感知注入新范式,恶意内容仅在智能体交互时显现,对用户不可见。为适应移动端界面约束和一次性交互场景,提出HG-IDA*,一种高效的单次优化方法,用于构造可绕过LVLM安全过滤器的越狱提示。实验表明,该方法能诱导未经授权的跨应用操作,在GPT-4o上实现82.5%的规划劫持率和75.0%的执行劫持率。研究揭示了移动智能体系统中此前被忽视的攻击面,强调需引入交互级信号进行防御。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vulnerable to visual prompt injections, stealthily executing such attacks without requiring system-level privileges remains challenging, as existing methods rely on persistent visual manipulations that are noticeable to users. We uncover a consistent discrepancy between human and agent interactions: automated agents generate near-zero contact touch signals. Building on this insight, we propose a new attack paradigm, agent-only perceptual injection, where malicious content is exposed only during agent interactions, while remaining not readily perceived by human users. To accommodate mobile UI constraints and one-shot interaction settings, we introduce HG-IDA*, an efficient one-shot optimization method for constructing jailbreak prompts that evade LVLM safety filters. Experiments demonstrate that our approach induces unauthorized cross-app actions, achieving 82.5% planning and 75.0% execution hijack rates on GPT-4o. Our findings highlight a previously underexplored attack surface in mobile agent systems and underscore the need for defenses that incorporate interaction-level signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。