通过多尺度渐进优化与自适应特征对齐,提升对多模态大模型的攻击转移性。
Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

- 分阶段从粗到细优化图像,结合中间层特征自适应选择与局部区域过滤。
- 在6个开源和6个闭源模型上,攻击成功率显著高于7种现有基线方法。
- 适合研究多模态模型安全性的研究人员,尤其关注对抗攻击与鲁棒性设计。
对抗扰动可使多模态大语言模型(MLLM)将正常图像误识别为特定目标物体,在自动驾驶、医疗诊断等安全关键场景中带来严重风险。因此,基于迁移的定向攻击对于理解并提升黑盒MLLM的鲁棒性至关重要。现有方法通常依赖替代编码器的最终全局特征,并锚定于原始分辨率的目标图像裁剪,导致其迁移能力与鲁棒性有限。为此,我们提出渐进式分辨率处理与自适应特征对齐(PRAF-Attack)框架,融合多尺度全局语义引导与稳健的中间层局部对齐机制。不同于仅对齐最终层的方法,我们设计了基于梯度一致性的自适应中间层选择机制,以识别跨替代编码器集合的可迁移层次特征,并采用自适应补丁级优化策略,通过高效补丁筛选保留高度相关局部区域。为克服对固定原始分辨率目标裁剪的依赖,我们提出渐进式分辨率处理策略,逐步从粗粒度到细粒度优化,使攻击能更好利用多尺度目标信息,实现更强迁移性。我们在包括六个开源模型和六个闭源商业API在内的多样化黑盒MLLM上评估PRAF-Attack,结果表明,相比七种最先进基线方法,本方法始终展现出更优的迁移性能。
原文摘要 · Abstract (English)
Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risks in safety-critical scenarios such as autonomous driving and medical diagnosis. This makes transfer-based targeted attacks crucial for understanding and improving black-box MLLM robustness. Existing transfer-based targeted attack methods typically rely on the final global features of the surrogate encoder and anchor optimization to original-resolution target crops, leading to their limited transferability and robustness. To address these challenges, we propose Progressive Resolution Processing and Adaptive Feature Alignment (PRAF-Attack), a targeted transfer-based attack framework that integrates multi-scale global semantic guidance with robust intermediate-layer local alignment. Unlike prior methods that align only the surrogate encoder's final layer, we design an adaptive feature alignment strategy that leverages intermediate representations to enhance transferability. Specifically, we introduce an adaptive intermediate layer selection mechanism to identify transferable hierarchical features across surrogate ensembles via gradient consistency, along with an adaptive patch-level optimization strategy that preserves highly correlated local regions through efficient patch filtering. To overcome the reliance on fixed original-resolution target crops, we propose a progressive resolution processing strategy that gradually refines optimization from coarse to fine, enabling the attack to better exploit target information at multiple scales and achieve stronger transferability. We evaluate PRAF-Attack on a diverse suite of black-box MLLMs, including six open-source models and six closed-source commercial APIs. Compared with seven state-of-the-art targeted attack baselines, the proposed PRAF-Attack consistently achieves superior transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。