arXiv:2608.09445cs.CRcs.CV2026-08

提出新方法防止扩散模型合并时继承后门,不依赖攻击细节也能有效检测并抑制可疑来源。

DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging

论文配图:DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging
图 1 · 摘自论文原文
  • 用无标签纯净数据和固定探测器评估各源块风险,动态调整贡献度。
  • 在14个源案例中10个实现最坏目标攻击成功率归零,剩余4个也无目标匹配。
  • 适合关注模型安全的开发者,尤其在未知攻击场景下仍能保持生成质量。

无条件扩散模型检查点合并假设源方可信,但受控的公开检查点可能在看似正常的生成中传递潜伏后门。由于无法获知受损源、触发条件或目标,且广泛净化可能降低图像质量,本文提出DiffSafeMerge(DSM):利用少量未标注的纯净数据集与固定的、攻击无关的应力探测器,对源块进行评分,将可疑贡献向可信参考值收缩,并在清洁去噪损失预算内选择衰减策略。我们评估了四种攻击、两个数据集及21种目标条件。在已有合并基础上,14个源案例中有10个实现最坏目标攻击成功率(ASR)为零;DSM不仅维持这些结果,在剩余四个案例中亦未出现目标匹配,包括三个基线ASR达48%–100%的情况。在两类数据集上均实现零最坏目标ASR的方法中,DSM在种子0对比中取得最低的平均FID值。

原文摘要 · Abstract (English)

Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean generation appears normal. Mitigation is difficult without knowing the compromised source, trigger, or target, and broad sanitization may degrade image quality. We introduce DiffSafeMerge (DSM), which uses a small unlabeled clean set and fixed, attack-agnostic stress probes to score source blocks, shrink suspicious contributions toward a trusted reference, and select attenuation under a clean denoising-loss budget. We evaluate four attacks, two datasets, and 21 target conditions. Intended merging already has zero worst-target ASR in 10 of 14 source cases; DSM preserves these outcomes and records no target match in the remaining four over three seeds, including three with baseline ASR of 48--100\%. Among methods with zero worst-target ASR on both datasets, DSM obtains the lowest case-averaged FID in the matched seed-0 comparison.

扩散模型模型安全后门防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。