arXiv:2502.14881cs.CRcs.CV2025-02综述被引 77

系统梳理大视觉语言模型安全问题,涵盖攻击、防御与评估方法。

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

  • 构建统一框架整合攻击、防御与评估三类研究
  • 对Deepseek Janus-Pro进行安全评测并分析结果
  • 适合关注多模态模型安全的研究者与开发者

随着大视觉语言模型(LVLMs)的快速发展,其安全性已成为关键研究方向。本综述全面分析了LVLM安全的核心方面,包括攻击、防御与评估方法。我们提出一个集成框架,将这些相互关联的组件统一起来,从生命周期视角区分训练与推理阶段,并进一步细分为子类别以深化理解。通过分析现有研究的局限性,我们指出了未来发展方向,旨在提升LVLM的鲁棒性。作为研究的一部分,我们对最新模型Deepseek Janus-Pro进行了安全评估,并提供理论分析。研究结果为推动LVLM安全发展、保障其在高风险真实场景中的可靠部署提供了策略建议。为促进该领域研究,我们维护了一个公开仓库,持续收集和更新最新的LVLM安全工作:https://github.com/XuankunRong/Awesome-LVLM-Safety。

原文摘要 · Abstract (English)

With the rapid advancement of Large Vision-Language Models (LVLMs), ensuring their safety has emerged as a crucial area of research. This survey provides a comprehensive analysis of LVLM safety, covering key aspects such as attacks, defenses, and evaluation methods. We introduce a unified framework that integrates these interrelated components, offering a holistic perspective on the vulnerabilities of LVLMs and the corresponding mitigation strategies. Through an analysis of the LVLM lifecycle, we introduce a classification framework that distinguishes between inference and training phases, with further subcategories to provide deeper insights. Furthermore, we highlight limitations in existing research and outline future directions aimed at strengthening the robustness of LVLMs. As part of our research, we conduct a set of safety evaluations on the latest LVLM, Deepseek Janus-Pro, and provide a theoretical analysis of the results. Our findings provide strategic recommendations for advancing LVLM safety and ensuring their secure and reliable deployment in high-stakes, real-world applications. This survey aims to serve as a cornerstone for future research, facilitating the development of models that not only push the boundaries of multimodal intelligence but also adhere to the highest standards of security and ethical integrity. Furthermore, to aid the growing research in this field, we have created a public repository to continuously compile and update the latest work on LVLM safety: https://github.com/XuankunRong/Awesome-LVLM-Safety .

大模型安全视觉语言模型攻击防御评测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。