arXiv:2509.23689cs.LG2025-09中稿 · The 29th Internati…被引 2

模型融合会提升对抗样本的迁移性,可能带来安全风险。

Merge Now, Regret Later: The Hidden Cost of Model Merging Is Adversarial Transferability

  • 通过8种融合方法、7个数据集验证对抗样本迁移性
  • 超过80%的攻击在融合后仍可成功迁移,防御失效
  • 越强的融合方法越易被迁移攻击,权重平均反而是最脆弱的

模型融合(MM)被视为多任务学习的有效替代方案,可在不访问训练数据的情况下将多个微调模型合并为一个跨任务保持性能的模型。尽管近期研究显示MM能增强对多种对抗攻击的鲁棒性,但对其在可迁移对抗样本攻击下影响的研究仍不足。本文系统评估了8种MM方法、7个数据集和6种攻击方法,共336种攻击设置。结果表明,MM无法可靠防御迁移攻击,迁移成功率超过80%。我们发现两大关键洞察:(1) 越强的融合方法反而越易受迁移攻击;(2) 减少表示偏差会增加迁移风险。值得注意的是,权重平均虽是最弱的融合方法,却是最易受攻击的。我们进一步分析其成因,说明低信息量攻击者亦可发起攻击,并提出潜在缓解方案。这些发现为安全敏感场景中部署模型融合提供了实用指导。

原文摘要 · Abstract (English)

Model Merging (MM) has proven to be an effective alternative to multi-task learning, where several fine-tuned models are merged, without access to the tasks' training data, into one model that retains performance across different tasks. Recent works have explored the security of MM, showing how MM can confer robustness against various adversarial attacks. However, none of them has sufficiently explored its impact on transfer attacks using transferable adversarial examples. In this work, we study the effect of MM on the transferability of adversarial examples. We perform comprehensive evaluations and statistical analysis consisting of eight MM methods, seven datasets, and six attack methods, sweeping over 336 distinct attack settings. Through it, we first challenge the prevailing notion of MM conferring free adversarial robustness, and show that MM cannot reliably defend against transfer attacks, with over 80% transfer rate. Moreover, we reveal two key insights for machine-learning practitioners regarding MM and transferability for a robust system design: (1) stronger MM methods increase vulnerability to transfer attacks and (2) mitigating representation bias increases vulnerability to transfer attacks. We also find weight averaging as an exception to (1), i.e., it is found to be the most vulnerable method to transfer attacks, despite being the weakest MM method. Finally, we analyze the underlying reasons for these findings, show how an adversary with limited information could launch an attack, and provide potential solutions. These findings offer actionable insights for deploying MM in security-sensitive machine-learning systems.

模型融合对抗攻击迁移性安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。