通过文本锚定的语义扰动,实现跨模型的多模态大模型越狱攻击
Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

- 在文本锚定的语义空间中优化可迁移扰动
- 在多个商用MLLM上实现高成功率越狱攻击
- 适用于研究模型安全与对抗防御的学者
多模态大语言模型(MLLMs)在视觉-语言交互中取得显著进展,但其安全对齐仍易受越狱攻击。关键挑战在于,文本空间中学习的安全行为无法可靠传递至跨模态融合表示,导致多模态输入可通过潜在语义线索被利用。本文提出文本锚定语义扰动攻击(TA-SPA),一种黑盒越狱框架,通过在文本锚定的语义空间中优化可迁移扰动。该方法结合文本锚定语义分解(TASF),促进跨模态语义因子与模态特异性残差的分离;以及语义保持增强(SPA),在保持语义一致性的同时多样化有害目标锚点。实验表明,该方法在多个商用MLLM上均具强攻击有效性,并在代表性防御下表现良好。额外控制与探测支持预期的语义因子分解,但不保证完全解耦,暗示应从表示层面而非仅输入层面进行安全对齐。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable perturbations in a text-anchored semantic space. TA-SPA integrates Text-Anchored Semantic Factorization (TASF), which encourages the separation of cross-modal semantic factors from modality-specific residuals, with Semantic-Preserving Augmentation (SPA), which diversifies harmful target anchors while preserving semantic consistency. Experiments show strong attack effectiveness and transfer to commercial MLLMs, with competitive performance under representative defenses. Additional controls and probing support the intended factorization without implying perfect disentanglement, motivating representation-level safety alignment beyond input-level filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。