通过频域正则化提升对抗攻击在闭源多模态大模型间的迁移能力。
Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs

- 在频域中用高通DCT抑制冗余全局结构,聚焦模型共享的视觉关注点。
- 引入无模型依赖的频域梯度正则,消除源模型特有高频噪声。
- 在15个主流闭源模型上验证,对GPT-5.4等模型表现最优。
多模态大语言模型(MLLMs)仍易受基于迁移的定向攻击影响,即在开源替代编码器上优化的扰动可泛化至闭源MLLMs。提升对抗迁移性的关键在于有效捕捉不同模型间共享的内在视觉关注点,使扰动对齐可迁移语义线索而非源模型特有行为。然而现有方法存在空间域特征冗余与源模型特有梯度信号问题,阻碍跨模型迁移。本文提出FRA-Attack,从统一的频域正则化视角解决上述挑战。针对特征对齐,采用对图像块特征的高通DCT目标,抑制冗余全局结构,将损失集中于携带模型内在视觉关注点的高频带;针对梯度优化,提出频域梯度正则化(FGR),一种无需源模型统计信息、仅依赖几何频率坐标的模型无关低通正则器,从而在不引入源模型特有高频伪影的同时保留可迁移的低频方向。二者结合形成统一的频域迁移增强机制。在7家厂商的15个旗舰级MLLM上进行大量实验表明,FRA-Attack在跨模型迁移性方面表现卓越,尤其在GPT-5.4、Claude-Opus-4.6和Gemini-3-flash上达到领先水平。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) remain vulnerable to transfer-based targeted attacks, where perturbations optimized on open-source surrogate encoders can generalize to closed-source MLLMs. A key challenge for improving adversarial transferability is to effectively capture the intrinsic visual focus shared across different models, such that perturbations align with transferable semantic cues rather than surrogate-specific behaviors. However, existing methods suffer from spatial-domain feature redundancy and surrogate-specific gradient signals, thereby hindering cross-model transferability. In this paper, we propose FRA-Attack, which addresses both challenges from a unified frequency-domain regularization perspective. For feature alignment, a high-pass DCT objective on patch features suppresses redundant global structures and concentrates the loss on the high-frequency band that carries the MLLMs' intrinsic visual focus. For gradient optimization, we introduce Frequency-domain Gradient Regularization (FGR), a \textit{model-agnostic} low-pass regularizer that modulates the surrogate gradient using only the geometric frequency coordinate, \textit{i.e.}, no surrogate-derived statistic is involved, so that FGR is model-agnostic by construction, removing surrogate-specific high-frequency artifacts while preserving transferable low-frequency directions. Together, the two components form a unified frequency-domain treatment of transferability. Extensive experiments on $15$ flagship MLLMs across $7$ vendors show that FRA-Attack achieves superior cross-model transferability, particularly with state-of-the-art performance on GPT-5.4, Claude-Opus-4.6 and Gemini-3-flash.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。