arXiv:2608.26506cs.LGcs.CL2026-08中稿 · EMNLP

模型合并会暴露新漏洞,只需一个后缀就能攻破多个合并模型。

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

论文配图:A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families
图 1 · 摘自论文原文
  • 通过优化合并空间生成可迁移的对抗性后缀
  • 在多种合并设置下成功率超80%,且防御无效
  • 适合研究模型安全与对抗攻击的学者

模型合并可在不额外训练的情况下融合多个微调模型,但其安全性影响尚未明确。以往研究主要归因于不安全的原始模型,假设单独对齐的模型合并后仍安全。然而,我们发现即使所有原始模型均安全对齐,基于同一预训练基础模型的合并仍存在未被重视的越狱风险。为此,我们提出一种新型攻击场景:攻击者在无法获取具体合并系数或原始模型权重的情况下,构造能跨共享基础模型的合并模型通用的越狱提示。为此,我们提出基域感知越狱(BAJ),将越狱生成建模为合并空间上的极小极大优化问题,以生成可跨合并模型家族迁移的对抗性后缀。在多种基础模型和合并设置下的实验表明,BAJ在不同条件下均保持超过80%的迁移成功率,且对现有防御手段具有鲁棒性。

原文摘要 · Abstract (English)

Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. In contrast, we show that model merging reveals a previously overlooked jailbreak risk rooted in the pretrained foundation model, even when all constituent models are individually safety-aligned. Motivated by this observation, we study a new threat setting where an attacker constructs jailbreak prompts that generalize across merged models sharing the same pretrained backbone, without access to the exact merging coefficients or constituent checkpoints. To exploit this phenomenon, we propose \textbf{Basin-Aware Jailbreak (BAJ)}, which formulates jailbreak generation as a min--max optimization over the merging space to produce transferable adversarial suffixes across merged model families. Experiments across diverse backbones and merging settings show that BAJ achieves consistently high transfer success rates and remains effective under existing defenses.

模型合并越狱攻击对抗样本安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。