通过分阶段概念对齐提升视觉语言模型安全性,防止恶意图像攻击。
PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment
- 引入安全概念瓶颈,分两阶段对齐视觉与语言模态的安全判断。
- 在多个基准测试中达到当前最佳安全防护效果,且通用性能影响小。
- 适合关注AI安全、可解释性及可控生成的研究者与开发者。
得益于大语言模型的强大能力,将预训练视觉编码器与大语言模型连接形成视觉语言模型(VLM)。然而,近期研究表明,VLM中的视觉模态极为脆弱,攻击者可通过视觉内容绕过大语言模型的安全对齐,实施有害攻击。为应对这一挑战,我们提出一种渐进式概念对齐策略PSA-VLM,通过引入安全模块作为概念瓶颈,增强视觉模态的安全对齐。通过使模型预测与特定安全概念对齐,提升了对风险图像的防御能力,在保持可解释性和可控性的同时,对整体性能影响极小。该方法采用两阶段训练:第一阶段计算开销低,但性能提升显著;第二阶段微调语言模型,进一步优化安全表现。在主流VLM安全基准上,本方法取得当前最优结果。
原文摘要 · Abstract (English)
Benefiting from the powerful capabilities of Large Language Models (LLMs), pre-trained visual encoder models connected to LLMs form Vision Language Models (VLMs). However, recent research shows that the visual modality in VLMs is highly vulnerable, allowing attackers to bypass safety alignment in LLMs through visually transmitted content, launching harmful attacks. To address this challenge, we propose a progressive concept-based alignment strategy, PSA-VLM, which incorporates safety modules as concept bottlenecks to enhance visual modality safety alignment. By aligning model predictions with specific safety concepts, we improve defenses against risky images, enhancing explainability and controllability while minimally impacting general performance. Our method is obtained through two-stage training. The low computational cost of the first stage brings very effective performance improvement, and the fine-tuning of the language model in the second stage further improves the safety performance. Our method achieves state-of-the-art results on popular VLM safety benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。