系统梳理大模型后门攻击威胁与防御进展
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges

- 梳理训练与推理阶段的后门攻击机制
- 总结检测与防御技术最新进展
- 适合关注AI安全的研究者与开发者
大型语言模型(LLMs)在网页搜索、医疗和软件开发等领域产生深远影响。然而,随着模型规模扩大,其面临日益严重的网络安全风险,尤其是后门攻击。攻击者可利用大模型强大的记忆能力,仅通过操纵少量训练数据即注入后门,当触发器出现时,模型会在下游应用中表现出恶意行为。此外,指令微调和人类反馈强化学习(RLHF)等新兴学习范式加剧了这一风险,因其依赖未完全受控的众包数据与人工反馈。本文全面综述了大模型在开发或推理过程中出现的新型后门威胁,涵盖近期在防御与检测策略方面的进展,并指出关键挑战,为未来研究指明方向。
原文摘要 · Abstract (English)
The advancement of Large Language Models (LLMs) has significantly impacted various domains, including Web search, healthcare, and software development. However, as these models scale, they become more vulnerable to cybersecurity risks, particularly backdoor attacks. By exploiting the potent memorization capacity of LLMs, adversaries can easily inject backdoors into LLMs by manipulating a small portion of training data, leading to malicious behaviors in downstream applications whenever the hidden backdoor is activated by the pre-defined triggers. Moreover, emerging learning paradigms like instruction tuning and reinforcement learning from human feedback (RLHF) exacerbate these risks as they rely heavily on crowdsourced data and human feedback, which are not fully controlled. In this paper, we present a comprehensive survey of emerging backdoor threats to LLMs that appear during LLM development or inference, and cover recent advancement in both defense and detection strategies for mitigating backdoor threats to LLMs. We also outline key challenges in addressing these threats, highlighting areas for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。