用隐蔽图像诱导网页智能体执行恶意操作,突破真实场景攻击限制。
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

- 基于扩散模型生成视觉无害的对抗性图像,仅在广告位等受限区域生效。
- 在SeeAct和OpenClaw上实现100%目标动作劫持成功率,攻击不可见且无明显异常。
- 适合安全研究人员评估多模态智能体漏洞,尤其关注真实业务场景威胁。
基于多模态大语言模型(MLLM)的网页智能体为可视化浏览器自动化提供了高效高精度解决方案;然而其天然扩大了攻击面,引入新型视觉漏洞。现有对抗评估多依赖宽松威胁模型和明显可见的伪造痕迹。本文研究一种受限漏洞检测场景:评估者作为未授权第三方(如商家或广告商),仅能控制语义合法、空间受限的区域(如广告位、赞助卡片或局部组件)。在此现实约束下,提出MIRAGE——一种针对目标下一步动作劫持的视觉间接提示注入框架。该方法利用扩散模型生成严格限定于攻击者可控边界的感知无害对抗图像。为在严苛条件下最大化攻击效果,引入结合曲率感知对抗扩散引导与稀疏暗像素残差扰动的鲁棒优化技术。对主流MLLM网页智能体框架SeeAct和OpenClaw的全面评估表明,MIRAGE具有高度有效性、真实性和隐蔽性。
原文摘要 · Abstract (English)
Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inherently expand the attack surface, introducing novel vision-based vulnerabilities. Existing adversarial evaluations targeting these agents frequently rely on permissive threat models and visually conspicuous artifacts. In this paper, we investigate a constrained vulnerability detection setting: a trusted web platform where the evaluator acts solely as an unprivileged third party, such as a merchant or advertiser, controlling only a semantically legitimate, spatially constrained region, such as an ad slot, a sponsored card, or a localized widget. Operating under these realistic constraints, we propose MIRAGE, a novel visual indirect prompt injection framework for targeted next-action hijacking. Our approach leverages diffusion models to generate perceptually benign adversarial images strictly confined to the attacker-controlled boundaries permitted by the trusted service provider. To maximize attack efficacy within such a restrictive setting, we introduce a robust optimization technique combining curvature-aware adversarial diffusion guidance with sparse, dark-pixel residual perturbations. Comprehensive evaluations against prominent MLLM web agent frameworks, specifically SeeAct and OpenClaw, empirically demonstrate the potency, realism, and stealth of our proposed MIRAGE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。