提出FORCE方法,提升视觉越狱攻击在不同模型间的迁移能力。
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- 通过纠正特征依赖问题,引导攻击探索更广的可行区域。
- 在多个闭源模型上实现高效越狱,成功率显著提升。
- 适合安全评估人员用于检测多模态大模型漏洞。
多模态大语言模型(MLLMs)融合新模态虽增强能力,但也引入新漏洞。简单视觉越狱攻击比复杂文本攻击更易操控开源模型,但跨模型迁移能力差,难以有效检测闭源模型漏洞。本文分析攻击损失景观,发现攻击集中于高尖锐区域,对参数微小变化敏感。进一步分析中间层特征与频域表示,揭示其过度依赖狭窄层特征和语义贫乏的频率成分。为此,提出特征过依赖修正(FORCE)方法,引导攻击探索更广的层特征可行区域,并根据语义内容重标频域特征影响。该方法消除层与频域特征的非泛化依赖,发现平坦化的可行区域,显著提升跨模型迁移能力。大量实验表明,该方法可有效支持对闭源MLLMs的视觉红队测试。
原文摘要 · Abstract (English)
The integration of new modalities enhances the capabilities of multimodal large language models (MLLMs) but also introduces additional vulnerabilities. In particular, simple visual jailbreaking attacks can manipulate open-source MLLMs more readily than sophisticated textual attacks. However, these underdeveloped attacks exhibit extremely limited cross-model transferability, failing to reliably identify vulnerabilities in closed-source MLLMs. In this work, we analyse the loss landscape of these jailbreaking attacks and find that the generated attacks tend to reside in high-sharpness regions, whose effectiveness is highly sensitive to even minor parameter changes during transfer. To further explain the high-sharpness localisations, we analyse their feature representations in both the intermediate layers and the spectral domain, revealing an improper reliance on narrow layer representations and semantically poor frequency components. Building on this, we propose a Feature Over-Reliance CorrEction (FORCE) method, which guides the attack to explore broader feasible regions across layer features and rescales the influence of frequency features according to their semantic content. By eliminating non-generalizable reliance on both layer and spectral features, our method discovers flattened feasible regions for visual jailbreaking attacks, thereby improving cross-model transferability. Extensive experiments demonstrate that our approach effectively facilitates visual red-teaming evaluations against closed-source MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。