arXiv:2502.05224cs.CRcs.AI2025-02综述被引 38

系统梳理大模型后门攻击与防御,助力构建更安全的AI应用

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations

  • 按训练时白盒攻击分类,梳理主流攻击方法
  • 综述现有防御手段,揭示攻防演进规律
  • 适合关注大模型安全的研究者和开发者

大型语言模型(LLMs)在理解与生成自然语言方面取得了显著进展,近年来广泛应用在医疗、金融、教育等多个领域。随着其能力提升和部署广泛,安全风险日益凸显。近年来,后门攻击技术随防御机制的完善及模型功能增强而持续演化。本文针对训练时白盒后门攻击,建立通用分类体系,系统梳理攻击方法与对应防御策略。通过全面总结现有研究成果,旨在为未来拓展攻击场景、构建更强防御机制提供参考,推动更鲁棒的大型语言模型发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved significantly advanced capabilities in understanding and generating human language text, which have gained increasing popularity over recent years. Apart from their state-of-the-art natural language processing (NLP) performance, considering their widespread usage in many industries, including medicine, finance, education, etc., security concerns over their usage grow simultaneously. In recent years, the evolution of backdoor attacks has progressed with the advancement of defense mechanisms against them and more well-developed features in the LLMs. In this paper, we adapt the general taxonomy for classifying machine learning attacks on one of the subdivisions - training-time white-box backdoor attacks. Besides systematically classifying attack methods, we also consider the corresponding defense methods against backdoor attacks. By providing an extensive summary of existing works, we hope this survey can serve as a guideline for inspiring future research that further extends the attack scenarios and creates a stronger defense against them for more robust LLMs.

大模型安全后门攻击防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。