arXiv:2502.05206cs.CRcs.AI2025-02综述被引 65

系统梳理大模型安全威胁与防御,助力AI可靠落地

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

  • 构建大模型安全威胁分类体系,覆盖攻击类型与机制
  • 归纳主流防御策略与评估基准,推动研究标准化
  • 强调跨领域协作与可持续安全实践的必要性

大模型凭借大规模预训练展现出强大学习与泛化能力,已成为对话系统、推荐、自动驾驶、内容生成、医疗诊断和科学发现等领域的核心。但其广泛应用也带来显著安全风险,涉及鲁棒性、可靠性与伦理问题。本综述系统梳理了视觉基础模型(VFMs)、大语言模型(LLMs)、视觉语言预训练(VLP)模型、视觉语言模型(VLMs)、扩散模型(DMs)及大模型驱动的智能体的安全研究。涵盖对抗攻击、数据投毒、后门攻击、越狱与提示注入、能耗延迟攻击、数据与模型提取攻击及新兴的智能体特异性威胁。总结各类攻击的防御策略,梳理常用数据集与评测基准。进一步指出当前挑战:亟需全面评估体系、可扩展有效的防御机制与可持续的数据实践。强调研究界与国际协作对构建综合防御系统的重要性。

原文摘要 · Abstract (English)

The rapid advancement of large models, driven by their exceptional abilities in learning and generalization through large-scale pre-training, has reshaped the landscape of Artificial Intelligence (AI). These models are now foundational to a wide range of applications, including conversational AI, recommendation systems, autonomous driving, content generation, medical diagnostics, and scientific discovery. However, their widespread deployment also exposes them to significant safety risks, raising concerns about robustness, reliability, and ethical implications. This survey provides a systematic review of current safety research on large models, covering Vision Foundation Models (VFMs), Large Language Models (LLMs), Vision-Language Pre-training (VLP) models, Vision-Language Models (VLMs), Diffusion Models (DMs), and large-model-powered Agents. Our contributions are summarized as follows: (1) We present a comprehensive taxonomy of safety threats to these models, including adversarial attacks, data poisoning, backdoor attacks, jailbreak and prompt injection attacks, energy-latency attacks, data and model extraction attacks, and emerging agent-specific threats. (2) We review defense strategies proposed for each type of attacks if available and summarize the commonly used datasets and benchmarks for safety research. (3) Building on this, we identify and discuss the open challenges in large model safety, emphasizing the need for comprehensive safety evaluations, scalable and effective defense mechanisms, and sustainable data practices. More importantly, we highlight the necessity of collective efforts from the research community and international collaboration. Our work can serve as a useful reference for researchers and practitioners, fostering the ongoing development of comprehensive defense systems and platforms to safeguard AI models.

大模型安全智能体安全防御机制综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。