系统梳理大模型在训练、推理和可用性阶段的攻击手段与防御策略。
A Survey of Attacks on Large Language Models
- 按训练、推理、可用性三阶段分类攻击方法
- 涵盖恶意滥用、隐私泄露等关键风险场景
- 适合关注大模型安全的开发者与研究人员
大型语言模型(LLMs)及其基于代理的应用已广泛部署于医疗诊断、金融分析、客户支持、机器人和自动驾驶等领域,显著提升了自然语言理解、推理与生成能力。然而,其广泛应用也暴露了严重的安全与可靠性风险,如恶意滥用、隐私泄露和服务中断,削弱用户信任并威胁社会安全。本文系统综述针对大语言模型及基于代理的攻击,按三个阶段组织:训练阶段攻击、推理阶段攻击、可用性与完整性攻击。对每类阶段中代表性且近期提出的攻击方法及其对应防御措施进行分析。希望本综述能为大模型安全提供全面理解与教学参考,提升对广泛部署的大模型应用固有风险的关注,并强调应对不断演进威胁的鲁棒缓解策略的紧迫性。
原文摘要 · Abstract (English)
Large language models (LLMs) and LLM-based agents have been widely deployed in a wide range of applications in the real world, including healthcare diagnostics, financial analysis, customer support, robotics, and autonomous driving, expanding their powerful capability of understanding, reasoning, and generating natural languages. However, the wide deployment of LLM-based applications exposes critical security and reliability risks, such as the potential for malicious misuse, privacy leakage, and service disruption that weaken user trust and undermine societal safety. This paper provides a systematic overview of the details of adversarial attacks targeting both LLMs and LLM-based agents. These attacks are organized into three phases in LLMs: Training-Phase Attacks, Inference-Phase Attacks, and Availability & Integrity Attacks. For each phase, we analyze the details of representative and recently introduced attack methods along with their corresponding defenses. We hope our survey will provide a good tutorial and a comprehensive understanding of LLM security, especially for attacks on LLMs. We desire to raise attention to the risks inherent in widely deployed LLM-based applications and highlight the urgent need for robust mitigation strategies for evolving threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。