arXiv:2501.18533cs.CVcs.CL2025-01被引 22

提升视觉语言模型安全推理能力,解决多图场景下的安全判断短板。

Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models

  • 构建多图像安全推理数据集,含细粒度安全思维链标注。
  • 在多图安全任务上显著降低攻击成功率,通用性能提升0.83%。
  • 适合需要高安全性的视觉模型应用,如医疗、自动驾驶。

大型视觉语言模型在众多任务中表现卓越,但在安全关键领域部署时仍面临挑战。现有安全微调方法主要针对文本或跨模态内容,难以应对复杂场景,且破坏帮助性与无害性之间的平衡。我们的评估揭示了安全推理能力缺失的瓶颈:模型缺乏安全视觉推理能力。为此,我们提出一种新数据集,将多图像输入与安全链式思考(CoT)标签结合,以细化推理逻辑。具体地,我们构建了面向多图像安全场景的指令跟随数据集 Multi-Image Safety (MIS),包含训练和测试集。实验表明,使用 MIS 微调 InternVL2.5-8B 在需安全视觉推理的复杂多图任务中,显著优于多个开源及API模型。该方法不仅实现优异的安全表现,还完全保留通用能力,未产生任何权衡。具体而言,微调后在五个通用基准上平均准确率提升0.83%,在多个安全基准上攻击成功率大幅下降。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) have achieved remarkable performance across a wide range of tasks. However, their deployment in safety-critical domains poses significant challenges. Existing safety fine-tuning methods, which focus on textual or multimodal content, fall short in addressing challenging cases or disrupt the balance between helpfulness and harmlessness. Our evaluation highlights a safety reasoning gap: these methods lack safety visual reasoning ability, leading to such bottlenecks. To address this limitation and enhance both visual perception and reasoning in safety-critical contexts, we propose a novel dataset that integrates multi-image inputs with safety Chain-of-Thought (CoT) labels as fine-grained reasoning logic to improve model performance. Specifically, we introduce the Multi-Image Safety (MIS) dataset, an instruction-following dataset tailored for multi-image safety scenarios, consisting of training and test splits. Our experiments demonstrate that fine-tuning InternVL2.5-8B with MIS significantly outperforms both powerful open-source models and API-based models in challenging multi-image tasks requiring safety-related visual reasoning. This approach not only delivers exceptional safety performance but also preserves general capabilities without any trade-offs. Specifically, fine-tuning with MIS increases average accuracy by 0.83% across five general benchmarks and reduces the Attack Success Rate (ASR) on multiple safety benchmarks by a large margin.

视觉语言模型安全推理多图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。