arXiv:2608.08542cs.LG2026-08

提出新基准,发现合并模型看似安全实则易被动态攻击突破。

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

论文配图:When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
图 1 · 摘自论文原文
  • 设计动态攻击测试框架,评估合并模型真实安全鲁棒性
  • 60%-76%的表面安全模型在语义攻击下被成功越狱
  • 适合关注模型融合安全性的研究者与实践者

模型合并已成为无需微调即可为对齐语言模型注入新技能的主流方法:从业者通过任务算术、TIES或DARE将数学、编程或领域专家的任务向量融入安全对齐的基础模型。这种便利性虽已知带来安全代价,但几乎所有证据均基于静态拒绝测试——固定有害提示的合规性评分。我们指出此法具有误导性,因安全对齐是‘浅层’的,集中于生成的前几项令牌,导致合并模型在静态拒绝上表现良好,却仍易受真实动态攻击影响。为此,我们提出SkillSafe-Bench,一个受控基准,以保守双评委‘与’规则评估技能合并模型的静态拒绝、自适应越狱鲁棒性及能力保留。在六种开源基础模型(五类,两尺度)上,静态安全性无法预测对抗鲁棒性:在语义模板攻击下,脆弱基模型(两个Qwen尺度和Gemma)有60%-76%被越狱,而其他模型(Llama、Phi-4)保持鲁棒。我们进一步揭示合并的静态影响具有基模型依赖性,通过无数据几何信号(任务向量与安全子空间重叠度)刻画同配方下的安全退化,并提出SubSafe-Merge,可投影消除该退化以保持能力。对于合并后的大型语言模型,自适应评估绝非可选项:最需要它的模型恰恰在静态筛查中表现安全。

原文摘要 · Abstract (English)

Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance. We argue this is misleading. Because safety alignment is "shallow," concentrated in the first few generated tokens, a merged model's static refusal can stay clean while a real adaptive attack still breaks it. We introduce SkillSafe-Bench, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule. Across six open-weight bases (five families, two scales), static safety does not predict robustness to attack: under a semantic template attack, safe-looking merges on the fragile bases (both Qwen scales and Gemma) are jailbroken 60-76% of the time while others (Llama, Phi-4) stay robust. We further show the static effect of merging is base-conditional, characterize same-recipe abliteration-style safety erosion through a data-free geometric signal (the overlap of a task vector with a safety subspace), and outline SubSafe-Merge, which projects this overlap away to remove that erosion at held capability. Adaptive evaluation is not optional for merged LLMs: the models that most need it look safe under static screening.

模型合并安全评测越狱攻击LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。