arXiv:2512.04895cs.AIcs.MA2025-12被引 1

用动态自适应攻击破解视觉提示隐藏漏洞,让模型被看不见的指令操控

Chameleon: Adaptive Adversarial Agents for Scaling-Based Visual Prompt Injection in Multimodal AI Systems

  • 设计可实时反馈优化的智能代理,动态生成抗缩放的恶意视觉提示
  • 在不同缩放比例下攻击成功率高达84.5%,远超传统静态攻击的32.1%
  • 适用于检测生产级多模态系统的安全缺陷,尤其关注自动化流程中的隐蔽威胁

多模态人工智能系统,尤其是视觉-语言模型(VLMs),已广泛应用于自动驾驶决策、文档自动化处理等关键场景。随着系统规模扩大,其依赖标准化预处理流程以高效处理多样输入。然而,对图像缩放等操作的依赖,带来了显著却常被忽视的安全漏洞。本应用于计算优化的缩放算法,可能被利用来隐藏人类无法察觉但能激活模型语义指令的恶意视觉提示。现有对抗攻击策略大多静态,未能适应现代智能体工作流的动态特性。为此,我们提出Chameleon——一种新型自适应对抗框架,旨在揭示并利用生产级VLM中的缩放漏洞。不同于传统静态攻击,Chameleon采用基于智能体的迭代优化机制,根据目标模型实时反馈动态调整图像扰动,生成能抵御标准缩放操作的鲁棒对抗样本,从而劫持下游任务执行。我们在Gemini 2.5 Flash模型上评估了该框架,实验表明,在不同缩放因子下,攻击成功率(ASR)达84.5%,显著高于静态基线攻击的32.1%。此外,攻击有效破坏了智能体工作流,使多步任务决策准确率下降超过45%。最后,我们讨论了这些漏洞的深远影响,并提出多尺度一致性检查作为必要防御机制。

原文摘要 · Abstract (English)

Multimodal Artificial Intelligence (AI) systems, particularly Vision-Language Models (VLMs), have become integral to critical applications ranging from autonomous decision-making to automated document processing. As these systems scale, they rely heavily on preprocessing pipelines to handle diverse inputs efficiently. However, this dependency on standard preprocessing operations, specifically image downscaling, creates a significant yet often overlooked security vulnerability. While intended for computational optimization, scaling algorithms can be exploited to conceal malicious visual prompts that are invisible to human observers but become active semantic instructions once processed by the model. Current adversarial strategies remain largely static, failing to account for the dynamic nature of modern agentic workflows. To address this gap, we propose Chameleon, a novel, adaptive adversarial framework designed to expose and exploit scaling vulnerabilities in production VLMs. Unlike traditional static attacks, Chameleon employs an iterative, agent-based optimization mechanism that dynamically refines image perturbations based on the target model's real-time feedback. This allows the framework to craft highly robust adversarial examples that survive standard downscaling operations to hijack downstream execution. We evaluate Chameleon against Gemini 2.5 Flash model. Our experiments demonstrate that Chameleon achieves an Attack Success Rate (ASR) of 84.5% across varying scaling factors, significantly outperforming static baseline attacks which average only 32.1%. Furthermore, we show that these attacks effectively compromise agentic pipelines, reducing decision-making accuracy by over 45% in multi-step tasks. Finally, we discuss the implications of these vulnerabilities and propose multi-scale consistency checks as a necessary defense mechanism.

对抗攻击视觉提示多模态安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。