arXiv:2501.02373cs.LGcs.CR2025-01中稿 · the ACM SIGSAC Con…被引 3

提出针对任务向量的新型后门攻击,可隐蔽破坏模型在多种操作下的表现。

BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors

  • 设计不对称复合后门,利用任务向量加减操作实现隐蔽植入
  • 在任务学习、遗忘和类比等场景下均实现接近100%攻击成功率
  • 现有防御手段均失效,揭示当前任务向量安全机制严重漏洞

大规模预训练模型中的任务算术技术通过任务向量(TVs)实现无需大量重训练的快速适配。用户可通过加减等简单运算进行模块化更新。然而,这种灵活性也带来了新的安全挑战。本文研究了任务向量在后门攻击下的脆弱性,揭示恶意方如何利用其破坏模型完整性。我们提出一种名为BadTV的新型后门攻击,该攻击通过设计不对称的复合后门,在任务学习、遗忘和类比操作中均保持高效。大量实验表明,BadTV在多种场景下均实现近乎完美的攻击成功率,对依赖任务算术的模型构成严重威胁。我们还评估了现有防御措施,发现它们无法有效检测或缓解该攻击。结果凸显了在实际部署中建立鲁棒防护机制的紧迫性。

原文摘要 · Abstract (English)

Task arithmetic in large-scale pre-trained models enables agile adaptation to diverse downstream tasks without extensive retraining. By leveraging task vectors (TVs), users can perform modular updates through simple arithmetic operations like addition and subtraction. Yet, this flexibility presents new security challenges. In this paper, we investigate how TVs are vulnerable to backdoor attacks, revealing how malicious actors can exploit them to compromise model integrity. By creating composite backdoors that are designed asymmetrically, we introduce BadTV, a backdoor attack specifically crafted to remain effective simultaneously under task learning, forgetting, and analogy operations. Extensive experiments show that BadTV achieves near-perfect attack success rates across diverse scenarios, posing a serious threat to models relying on task arithmetic. We also evaluate current defenses, finding they fail to detect or mitigate BadTV. Our results highlight the urgent need for robust countermeasures to secure TVs in real-world deployments.

后门攻击任务向量模型安全预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。