提出高效防御多模态大模型越狱攻击的新框架,显著提升安全性和泛化能力。
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
- 采用动态联合优化策略,自动调节正负样本训练权重
- 在五种攻击方法上平均提升34%防御效果,保持原有性能
- 适合需要高安全性的智能机器人等多模态系统应用
针对多模态大模型(MLLMs)在越狱攻击下的脆弱性,现有方法仍面临两大挑战:如何高效调整海量参数,以及如何在视觉与文本双模态下确保鲁棒性。为此,本文提出一种高效端到端对抗训练框架E$^2$AT,用于应对视觉与文本双重攻击。在视觉方面,E$^2$AT引入基于投影器的对抗训练模块,在特征层面对齐攻击样本;在训练目标上,提出动态联合多模态优化(DJMO)策略,通过动态调整正常与对抗目标的权重,增强对越狱攻击的泛化能力。在三个主流MLLM上,使用五种典型越狱攻击方法进行实验,结果表明E$^2$AT在文本与图像模态上平均优于现有基线34%,同时维持原始任务性能。真实世界具身智能系统的评估也验证了其实际应用价值。代码已公开。
原文摘要 · Abstract (English)
Research endeavors have been made in learning robust Multimodal Large Language Models (MLLMs) against jailbreak attacks. However, existing methods for improving MLLMs' robustness still face critical challenges: \ding{172} how to efficiently tune massive weight parameters and \ding{173} how to ensure robustness against attacks across both visual and textual modalities. To this end, we propose an \textbf{E}fficient \textbf{E}nd-to-end \textbf{A}dversarial \textbf{T}raining (E$^2$AT) framework for both visual and textual adversarial attacks. Specifically, for the visual aspect, E$^2$AT incorporates an efficient projector-based AT module that aligns the attack samples at the feature level. For training objectives, we propose a Dynamic Joint Multimodal Optimization (DJMO) strategy to enhance generalization ability against jailbreak attacks by dynamically adjusting weights between normal and adversarial objectives. Extensive experiments are conducted with five major jailbreak attack methods across three mainstream MLLMs. Results demonstrate that our E$^2$AT achieves the state-of-the-art performance, outperforming existing baselines by an average margin of 34\% across text and image modalities, while maintaining clean task performance. Furthermore, evaluations of real-world embodied intelligent systems highlight the practical applicability of E$^2$AT, paving the way for the development of more secure and reliable multimodal systems. Our code is available on \href{https://anonymous.4open.science/r/E2AT_568}{\textcolor{red}{https://anonymous.4open.science/r/E2AT\_568}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。