攻击多模态智能体,用图文混合干扰让其执行恶意指令
Manipulating Multimodal Agents via Cross-Modal Prompt Injection

- 通过视觉嵌入与文本引导协同设计对抗性扰动
- 在多种任务中攻击成功率提升超30.1%
- 适合关注AI安全与防御的研究者阅读
多模态大语言模型通过融合语言与视觉模态及外部数据源,重塑了智能体范式,使其能更好理解人类指令并执行复杂任务。然而本文揭示了一类此前被忽视的关键安全漏洞:跨模态提示注入攻击。为此,我们提出CrossInject攻击框架,通过在多个模态中嵌入对抗性扰动,使其与目标恶意内容对齐,从而劫持智能体决策过程并执行未授权任务。该方法包含两个关键组件:首先引入视觉潜空间对齐,基于文本生成图像模型优化对抗特征,使图像隐含恶意任务执行线索;其次提出文本引导增强,利用大语言模型通过对抗元提示构建黑盒防御系统提示,并生成可诱导智能体高配合度的恶意文本指令。大量实验表明,该方法优于现有最先进攻击,在多种任务中攻击成功率至少提升+30.1%。此外,我们在真实世界多模态自主智能体上验证了攻击有效性,凸显其对安全关键应用的潜在影响。
原文摘要 · Abstract (English)
The emergence of multimodal large language models has redefined the agent paradigm by integrating language and vision modalities with external data sources, enabling agents to better interpret human instructions and execute increasingly complex tasks. However, in this paper, we identify a critical yet previously overlooked security vulnerability in multimodal agents: cross-modal prompt injection attacks. To exploit this vulnerability, we propose CrossInject, a novel attack framework in which attackers embed adversarial perturbations across multiple modalities to align with target malicious content, allowing external instructions to hijack the agent's decision-making process and execute unauthorized tasks. Our approach incorporates two key coordinated components. First, we introduce Visual Latent Alignment, where we optimize adversarial features to the malicious instructions in the visual embedding space based on a text-to-image generative model, ensuring that adversarial images subtly encode cues for malicious task execution. Subsequently, we present Textual Guidance Enhancement, where a large language model is leveraged to construct the black-box defensive system prompt through adversarial meta prompting and generate an malicious textual command that steers the agent's output toward better compliance with attackers' requests. Extensive experiments demonstrate that our method outperforms state-of-the-art attacks, achieving at least a +30.1% increase in attack success rates across diverse tasks. Furthermore, we validate our attack's effectiveness in real-world multimodal autonomous agents, highlighting its potential implications for safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。