arXiv:2605.28116cs.CRcs.AI2026-05

用用户生成内容伪装攻击,让手机AI助手误操作

MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

论文配图:MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
图 1 · 摘自论文原文
  • 将普通截图转为攻击样本,注入隐蔽恶意文本
  • 在10个应用上测试,攻击成功率23%-30%
  • 攻击样本更逼真,人眼难辨,防御需新思路

由视觉语言模型驱动的移动GUI代理将屏幕视为像素图像,难以区分可信界面元素与用户生成内容。我们提出MIRAGE(移动真实对抗性界面示例注入),通过在普通用户生成内容区域插入攻击者控制的文本,将良性移动端截图转化为提示注入样本,无需修改代理、应用或操作系统。该方法分三阶段:定位器识别可操控区域,生成器合成上下文相关载荷并以原生样式渲染,策展者调节真实感并平衡样本分布。关键挑战在于保持样本与真实用户内容视觉一致的同时实现误导;通过分离覆盖范围、真实感与分布均衡三个阶段解决。在涵盖十款应用和十一类攻击意图的1,111样本基准上,所有五款评估的VLM代理均易受攻击,成功率达23%-30%。MIRAGE的人类真实感评分(3.02/5)高于最强现有攻击(2.52/5)。进一步发现,单样本真实感与攻击成功率无相关性,仅靠视觉质量过滤无法有效防御此威胁。

原文摘要 · Abstract (English)

Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot reliably separate trusted interface elements from user-generated content. We present MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples), a pipeline that turns benign mobile screenshots into prompt-injection samples by placing attacker-controlled text into ordinary user-generated content regions, without modifying the agent, the application, or the operating system. MIRAGE operates in three stages: a Localizer identifies user-controllable regions on the screenshot, a Generator synthesises context-aware payloads and renders them in the application's native style, and a Curator moderates realism and balances the samples across applications, region types, and attack intents. A key challenge is that an injected screenshot must stay visually indistinguishable from genuine user content while still diverting the agent; we address this by separating the stages that control reach, realism, and distributional balance. On a 1,111-sample benchmark spanning ten applications and eleven attack intents, all five evaluated VLM agents are vulnerable, with attack success rates of 23%-30%, and MIRAGE scores higher on human realism ratings than the strongest prior attack (3.02 versus 2.52 out of 5). We further find that per-sample realism and attack success are uncorrelated, so visual-quality filtering alone cannot reliably defend against this threat.

AI安全移动攻击视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。