首份覆盖LLM全生命周期安全的综述,从数据到部署全面梳理风险与方向。
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
- 构建包含数据、预训练、后训练、部署到商业化的全栈安全框架
- 基于800+论文系统梳理各阶段安全问题,提供完整知识图谱
- 聚焦数据生成、对齐技术、模型编辑等前沿方向,适合安全研究者参考
大型语言模型(LLMs)的显著成功为实现人工智能通用性开辟了新路径,其在各类应用中表现出前所未有的性能。随着LLMs在科研和产业领域日益重要,其安全与风险问题愈发受到学术界、企业及各国政府关注。现有安全综述多集中于生命周期的单一环节,如部署或微调阶段,缺乏对整个LLM“生命链”的系统理解。为此,本文首次提出“全栈”安全概念,系统性地涵盖从数据准备、预训练、后训练、部署到最终商业化全过程的安全挑战。相比现有综述,本工作具有三大优势:(I) 全面视角:将完整生命周期定义为数据准备、预训练、后训练、部署与商业化;(II) 广泛文献支撑:基于超过800篇论文的系统性回顾,确保覆盖全面、组织有序;(III) 独特洞见:通过系统分析,为各章节提炼可靠路线图与展望,识别出数据生成安全、对齐技术、模型编辑及基于LLM的智能体系统等关键研究方向,为后续研究提供有力指引。
原文摘要 · Abstract (English)
The remarkable success of Large Language Models (LLMs) has illuminated a promising pathway toward achieving Artificial General Intelligence for both academic and industrial communities, owing to their unprecedented performance across various applications. As LLMs continue to gain prominence in both research and commercial domains, their security and safety implications have become a growing concern, not only for researchers and corporations but also for every nation. Currently, existing surveys on LLM safety primarily focus on specific stages of the LLM lifecycle, e.g., deployment phase or fine-tuning phase, lacking a comprehensive understanding of the entire "lifechain" of LLMs. To address this gap, this paper introduces, for the first time, the concept of "full-stack" safety to systematically consider safety issues throughout the entire process of LLM training, deployment, and eventual commercialization. Compared to the off-the-shelf LLM safety surveys, our work demonstrates several distinctive advantages: (I) Comprehensive Perspective. We define the complete LLM lifecycle as encompassing data preparation, pre-training, post-training, deployment and final commercialization. To our knowledge, this represents the first safety survey to encompass the entire lifecycle of LLMs. (II) Extensive Literature Support. Our research is grounded in an exhaustive review of over 800+ papers, ensuring comprehensive coverage and systematic organization of security issues within a more holistic understanding. (III) Unique Insights. Through systematic literature analysis, we have developed reliable roadmaps and perspectives for each chapter. Our work identifies promising research directions, including safety in data generation, alignment techniques, model editing, and LLM-based agent systems. These insights provide valuable guidance for researchers pursuing future work in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。