构建全面安全评测集并设计可插拔防护模块,提升视觉语言模型安全性。
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
- 构建涵盖五类图文组合的全维度安全评测基准HoliSafe-Bench。
- 引入视觉警卫模块VGM,使模型能解释拒绝理由并提升安全性能。
- 模块化设计支持多种大模型接入,适用于安全增强研究与实践。
尽管已有研究致力于提升视觉语言模型(VLMs)的安全性,现有方法仍存在两大缺陷:一是安全调优数据集和评测基准仅部分考虑图像-文本交互导致的有害内容,常忽略看似无害组合带来的上下文不安全输出,导致模型在未见配置下易受越狱攻击;二是现有方法主要依赖数据驱动调优,缺乏架构层面的创新以内在强化安全性。为此,本文提出综合性安全数据集与评测基准HoliSafe,覆盖全部五种安全/不安全图像-文本组合,为训练与评估提供更稳健基础(HoliSafe-Bench)。进一步提出一种新型模块化框架,引入视觉警卫模块(VGM),用于评估输入图像对VLM的潜在危害性。该模块使模型兼具双重功能:不仅学习生成更安全响应,还能提供可解释的危害性分类以说明拒绝决策。其关键优势在于模块化设计,作为即插即用组件,可无缝集成于不同规模的预训练VLM中。实验表明,基于HoliSafe训练的Safe-VLM结合VGM,在多个VLM评测基准上达到当前最优安全性能。此外,HoliSafe-Bench揭示了现有VLM模型的关键漏洞。我们期望HoliSafe与VGM能推动更具鲁棒性和可解释性的多模态安全研究,拓展未来多模态对齐的新方向。
原文摘要 · Abstract (English)
Despite emerging efforts to enhance the safety of Vision-Language Models (VLMs), current approaches face two main shortcomings. 1) Existing safety-tuning datasets and benchmarks only partially consider how image-text interactions can yield harmful content, often overlooking contextually unsafe outcomes from seemingly benign pairs. This narrow coverage leaves VLMs vulnerable to jailbreak attacks in unseen configurations. 2) Prior methods rely primarily on data-centric tuning, with limited architectural innovations to intrinsically strengthen safety. We address these gaps by introducing a holistic safety dataset and benchmark, \textbf{HoliSafe}, that spans all five safe/unsafe image-text combinations, providing a more robust basis for both training and evaluation (HoliSafe-Bench). We further propose a novel modular framework for enhancing VLM safety with a visual guard module (VGM) designed to assess the harmfulness of input images for VLMs. This module endows VLMs with a dual functionality: they not only learn to generate safer responses but can also provide an interpretable harmfulness classification to justify their refusal decisions. A significant advantage of this approach is its modularity; the VGM is designed as a plug-in component, allowing for seamless integration with diverse pre-trained VLMs across various scales. Experiments show that Safe-VLM with VGM, trained on our HoliSafe, achieves state-of-the-art safety performance across multiple VLM benchmarks. Additionally, the HoliSafe-Bench itself reveals critical vulnerabilities in existing VLM models. We hope that HoliSafe and VGM will spur further research into robust and interpretable VLM safety, expanding future avenues for multimodal alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。