通过动态演化对抗攻击提升多模态模型安全对齐效果
Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution
- 构建可进化的对抗生成器,自动分解并演化恶意指令结构
- 在红队测试中使越狱攻击成功率显著提升,防御方更抗干扰
- 适合关注多模态安全与持续对抗训练的研究者和工程师
对抗行为在对齐大语言模型与人类价值观中起核心作用。然而,现有对齐方法多依赖静态对抗设置,严重限制了鲁棒性,尤其在攻击面更大的多模态场景下。本文提出一种超越静态对抗监督的共进化对齐框架——CEMMA(Co-Evolutionary Multi-Modal Alignment),实现自动化、自适应的多模态安全对齐。我们设计了进化攻击者,将对抗提示分解为方法模板与有害意图,并通过变异、交叉与差分进化等遗传算子,使简单初始攻击能继承复杂越狱技巧的结构优势。同时,自适应防御者在合成的困难负样本上迭代更新,形成闭环适应过程。实验表明,进化攻击者显著提高红队越狱攻击成功率(ASR),而自适应防御者在多个基准上增强鲁棒性与泛化能力,数据效率更高,且不引发过度良性拒绝,兼容推理时防御机制如AdaShield。
原文摘要 · Abstract (English)
Adversarial behavior plays a central role in aligning large language models with human values. However, existing alignment methods largely rely on static adversarial settings, which fundamentally limit robustness, particularly in multimodal settings with a larger attack surface. In this work, we move beyond static adversarial supervision and introduce co-evolutionary alignment with evolving attacks, instantiated by CEMMA (Co-Evolutionary Multi-Modal Alignment), an automated and adaptive framework for multimodal safety alignment. We introduce an Evolutionary Attacker that decomposes adversarial prompts into method templates and harmful intents. By employing genetic operators, including mutation, crossover, and differential evolution, it enables simple seed attacks to inherit the structural efficacy of sophisticated jailbreaks. The Adaptive Defender is iteratively updated on the synthesized hard negatives, forming a closed-loop process that adapts alignment to evolving attacks. Experiments show that the Evolutionary Attacker substantially increases red-teaming jailbreak attack success rate (ASR), while the Adaptive Defender improves robustness and generalization across benchmarks with higher data efficiency, without inducing excessive benign refusal, and remains compatible with inference-time defenses such as AdaShield.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。