系统梳理大模型全生命周期安全风险,揭示攻击链与防御难点。
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
- 按数据收集到部署共8个阶段分析漏洞,聚焦信任边界失效
- 指出模型错误因权限委派被放大,单一防护难组合生效
- 适合关注大模型安全的开发者、安全研究者和企业架构师
大语言模型已从单纯文本生成扩展至检索管道、企业助手、编程环境、机器人系统及安全运维流程中,能读取私有数据、调用工具、写入文件、执行代码并跨组织行动。安全风险不再仅来自模型权重,而是贯穿数据、提示、输出、工具、记忆与用户权限的完整生命周期与应用栈。本文基于生命周期与应用栈视角,系统梳理大模型系统的漏洞文献,涵盖数据收集、预训练、后训练对齐、模型封装与供应链、检索与记忆、提示与推理、工具/代理执行、部署与维护共八个阶段。每个阶段分析攻击者能力、受影响安全目标、典型攻击、实际风险、评估方法与防御策略。进一步将大模型特有漏洞映射至机密性、完整性、可用性、安全性、隐私、公平性、问责制与控制权等目标。区别于仅列攻击名称的分类法,本系统强调信任边界崩溃、不可信数据变可执行指令、权限委派放大模型错误,以及点状防御难以组合等问题。最后提出安全大模型系统的七大研究方向:组合安全、溯源感知检索、工具调用约束、长周期代理评估、隐私保护适应、真实红队测试与部署级应急响应。
原文摘要 · Abstract (English)
Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous agents that can read private data, call tools, write files, execute code, and act across organizational boundaries. This shift changes the security problem: risks do not arise from the model weights alone, but from the full lifecycle and application stack through which data, prompts, model outputs, tools, memories, and user authority interact. This paper systematizes the literature on vulnerabilities in large language model systems through a lifecycle and application-stack lens. We organize attacks across eight stages: data collection, pretraining, post-training alignment, model packaging and supply chain, retrieval and memory, prompting and inference, tool/agent execution, and deployment/maintenance. For each stage, we analyze attacker capabilities, affected security objectives, representative attacks, practical risks, evaluation practices, and defenses. We further map LLM-specific vulnerabilities to confidentiality, integrity, availability, safety, privacy, fairness, accountability, and agency-control objectives. Unlike taxonomies that list isolated attack names, the proposed systematization emphasizes where trust boundaries fail, how untrusted data becomes executable instruction, how delegated authority amplifies model errors, and why point defenses rarely compose. We close with a research agenda for secure LLM systems, including compositional security, provenance-aware retrieval, tool-call containment, long-horizon agent evaluation, privacy-preserving adaptation, realistic red teaming, and deployment-grade incident response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。