arXiv:2505.00976cs.CRcs.AI2025-05综述被引 27

系统梳理大模型攻防技术,揭示安全挑战与应对路径

Attack and defense techniques in large language models: A survey and new perspectives

  • 分类梳理对抗提示、模型窃取等攻击手段及其机制
  • 提出预防与检测两类防御策略,指出当前局限性
  • 强调可解释安全与标准化评估等关键开放问题

大型语言模型(LLMs)在自然语言处理任务中占据核心地位,但其漏洞带来了重大安全与伦理挑战。本文系统综述了LLMs攻防技术的演进态势,将攻击分为对抗提示攻击、优化攻击、模型窃取以及针对应用层面的攻击,详细阐述其作用机制与影响。进而分析了以预防和检测为基础的防御策略。尽管已有进展,但面对动态威胁环境,仍需平衡可用性与鲁棒性,并解决防御实施中的资源约束。文章指出若干开放问题:亟需自适应可扩展的防御机制、可解释的安全技术以及标准化评估框架。本综述为构建安全可靠的LLMs提供行动洞察与研究方向,强调跨学科协作与伦理考量对缓解真实应用场景风险的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become central to numerous natural language processing tasks, but their vulnerabilities present significant security and ethical challenges. This systematic survey explores the evolving landscape of attack and defense techniques in LLMs. We classify attacks into adversarial prompt attack, optimized attacks, model theft, as well as attacks on application of LLMs, detailing their mechanisms and implications. Consequently, we analyze defense strategies, including prevention-based and detection-based defense methods. Although advances have been made, challenges remain to adapt to the dynamic threat landscape, balance usability with robustness, and address resource constraints in defense implementation. We highlight open problems, including the need for adaptive scalable defenses, explainable security techniques, and standardized evaluation frameworks. This survey provides actionable insights and directions for developing secure and resilient LLMs, emphasizing the importance of interdisciplinary collaboration and ethical considerations to mitigate risks in real-world applications.

大模型安全攻防技术系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。