梳理大模型安全威胁,揭示攻击与防御的最新挑战
Security Concerns for Large Language Models: A Survey
- 按攻击阶段分类:提示劫持、训练污染、恶意使用、自主代理风险
- 2022-2025年学术工业研究综述,覆盖主流威胁与防御手段
- 聚焦自主代理安全,提出多层防护框架需求
大型语言模型(如ChatGPT)在自然语言处理领域引发革命,但其能力也带来新型安全漏洞。本文系统综述这些新兴威胁,将其归类为:推理阶段通过提示操纵实施的攻击;训练阶段攻击;恶意使用者的滥用;以及自主语言模型代理固有的风险。近年来,后一类问题受到越来越多关注。我们总结了2022至2025年间学术界与工业界的代表性研究,分析现有防御机制及其局限性,并指出保障基于大模型应用安全的开放挑战。最后强调,需发展稳健、多层次的安全策略,确保大模型安全且有益。
原文摘要 · Abstract (English)
Large Language Models (LLMs) such as ChatGPT and its competitors have caused a revolution in natural language processing, but their capabilities also introduce new security vulnerabilities. This survey provides a comprehensive overview of these emerging concerns, categorizing threats into several key areas: inference-time attacks via prompt manipulation; training-time attacks; misuse by malicious actors; and the inherent risks in autonomous LLM agents. Recently, a significant focus is increasingly being placed on the latter. We summarize recent academic and industrial studies from 2022 to 2025 that exemplify each threat, analyze existing defense mechanisms and their limitations, and identify open challenges in securing LLM-based applications. We conclude by emphasizing the importance of advancing robust, multi-layered security strategies to ensure LLMs are safe and beneficial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。