arXiv:2602.13427cs.CRcs.AI2026-02被引 1

研究大模型中隐藏的偏见后门,揭示攻击与防御的新弱点。

Backdooring Bias in Large Language Models

  • 白盒环境下测试语法与语义触发的后门攻击
  • 语义攻击更易制造负面偏见,但正向偏见难诱导
  • 现有防御手段会严重损害模型性能或计算成本高

大型语言模型在需引导特定话题偏见的场景中应用日益广泛,后门攻击可被用于生成此类模型。以往研究多聚焦黑盒攻击,针对模型构建者;但在偏见操纵场景中,模型构建者本身可能为攻击者,需采用白盒威胁模型,此时攻击者对训练数据的污染和操控能力显著增强。此外,尽管语义触发后门研究增多,多数工作仍局限于语法触发攻击。为此,本文基于超过1000次实验,在更高污染率与更强数据增强下,分析白盒环境中语法与语义触发后门攻击的潜力。同时考察两种代表性防御范式——模型内生与模型外在后门移除——的效能。结果发现:两类攻击均能有效诱导目标行为且保持较高模型效用,但语义触发攻击更擅长制造负面偏见,而两者均难以诱发正面偏见。尽管两类防御可缓解攻击,但均导致显著性能下降或高计算开销。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in settings where inducing a bias toward a certain topic can have significant consequences, and backdoor attacks can be used to produce such models. Prior work on backdoor attacks has largely focused on a black-box threat model, with an adversary targeting the model builder's LLM. However, in the bias manipulation setting, the model builder themselves could be the adversary, warranting a white-box threat model where the attacker's ability to poison, and manipulate the poisoned data is substantially increased. Furthermore, despite growing research in semantically-triggered backdoors, most studies have limited themselves to syntactically-triggered attacks. Motivated by these limitations, we conduct an analysis consisting of over 1000 evaluations using higher poisoning ratios and greater data augmentation to gain a better understanding of the potential of syntactically- and semantically-triggered backdoor attacks in a white-box setting. In addition, we study whether two representative defense paradigms, model-intrinsic and model-extrinsic backdoor removal, are able to mitigate these attacks. Our analysis reveals numerous new findings. We discover that while both syntactically- and semantically-triggered attacks can effectively induce the target behaviour, and largely preserve utility, semantically-triggered attacks are generally more effective in inducing negative biases, while both backdoor types struggle with causing positive biases. Furthermore, while both defense types are able to mitigate these backdoors, they either result in a substantial drop in utility, or require high computational overhead.

后门攻击模型偏见大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。