构建多模态安全推理数据集,提升视觉语言模型的抗攻击能力
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
- 基于安全政策设计多模态推理数据,支持跨模态细粒度判断
- 在文本和图文攻击下,模型鲁棒性显著提升,通用推理性能不降反升
- 适合关注AI安全、模型对齐的研究者与实践者
视觉语言模型(VLMs)在多模态推理任务中取得显著进展,但其增强的思维链能力也带来了新型安全风险,易受有害多模态提示诱导产生不当行为。现有安全对齐方法主要针对单模态语言模型,难以应对多模态输入带来的复杂威胁。当前安全数据集缺乏细粒度、以政策为依据的推理标注。本文提出 {MSR-Align},一个高质量的多模态安全推理数据集,支持在视觉与文本模态上基于标准化安全政策进行细粒度、深思熟虑的推理。数据生成流程强调多模态多样性、政策依据性推理,并采用强多模态判别器进行严格质量筛选。大量实验表明,在 MSR-Align 上微调的 VLM 能显著提升对文本及图文越狱攻击的鲁棒性,同时保持或增强通用推理能力。该数据集为推进具备推理能力的 VLM 安全对齐提供了可扩展的有效基础。数据集已公开于 https://huggingface.co/datasets/Leigest/MSR-Align。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have achieved remarkable progress in multimodal reasoning tasks through enhanced chain-of-thought capabilities. However, this advancement also introduces novel safety risks, as these models become increasingly vulnerable to harmful multimodal prompts that can trigger unethical or unsafe behaviors. Existing safety alignment approaches, primarily designed for unimodal language models, fall short in addressing the complex and nuanced threats posed by multimodal inputs. Moreover, current safety datasets lack the fine-grained, policy-grounded reasoning required to robustly align reasoning-capable VLMs. In this work, we introduce {MSR-Align}, a high-quality Multimodal Safety Reasoning dataset tailored to bridge this gap. MSR-Align supports fine-grained, deliberative reasoning over standardized safety policies across both vision and text modalities. Our data generation pipeline emphasizes multimodal diversity, policy-grounded reasoning, and rigorous quality filtering using strong multimodal judges. Extensive experiments demonstrate that fine-tuning VLMs on MSR-Align substantially improves robustness against both textual and vision-language jailbreak attacks, while preserving or enhancing general reasoning performance. MSR-Align provides a scalable and effective foundation for advancing the safety alignment of reasoning-capable VLMs. Our dataset is made publicly available at https://huggingface.co/datasets/Leigest/MSR-Align.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。