arXiv:2507.11112cs.CLcs.CR2025-07被引 3

多触发器攻击可让大模型同时藏多个后门,且彼此不干扰。

Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs

  • 提出多触发器共存框架,揭示后门交互机制。
  • 相似触发器在替换或分隔时仍能稳定激活,抗干扰强。
  • 基于权重差异的局部重训练法,低开销有效清除后门。

近期研究发现大型语言模型易受数据投毒攻击,恶意训练样本会嵌入特定输入模式触发的隐藏行为。然而,现有工作多聚焦单一触发词,对触发机制及多触发器交互理解有限。本文提出一种研究框架,证明多个不同后门触发器可在单个模型中并存且互不干扰,支持攻击者同时植入多个触发器。利用高嵌入相似性的多触发器,我们验证其在令牌被替换或长距离分隔时仍具鲁棒激活能力。结果揭示了更广泛持久的模型漏洞。为此,我们提出一种后处理恢复方法,通过分层权重差异分析选择性重训模型特定组件,仅需少量参数更新即可有效移除触发行为,为抵御多触发器投毒提供高效实用方案。

原文摘要 · Abstract (English)

Recent studies have shown that Large Language Models (LLMs) are vulnerable to data poisoning attacks, where malicious training examples embed hidden behaviours triggered by specific input patterns. However, most existing works assume a phrase and focus on the attack's effectiveness, offering limited understanding of trigger mechanisms and how multiple triggers interact within the model. In this paper, we present a framework for studying poisoning in LLMs. We show that multiple distinct backdoor triggers can coexist within a single model without interfering with each other, enabling adversaries to embed several triggers concurrently. Using multiple triggers with high embedding similarity, we demonstrate that poisoned triggers can achieve robust activation even when tokens are substituted or separated by long token spans. Our findings expose a broader and more persistent vulnerability surface in LLMs. To mitigate this threat, we propose a post hoc recovery method that selectively retrains specific model components based on a layer-wise weight difference analysis. Our method effectively removes the trigger behaviour with minimal parameter updates, presenting a practical and efficient defence against multi-trigger poisoning.

后门攻击大模型安全数据投毒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。