PRISM通过分阶段推理提升多模态模型安全,有效应对复杂威胁。
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
- 采用四阶段结构化推理流程,分类处理三类多模态安全问题。
- 在JailbreakV-28K和VLBreak上攻击成功率显著降低,对自适应攻击更鲁棒。
- 适合关注模型安全性与实用性平衡的研究者与开发者。
保障视觉语言模型(VLMs)的安全性是一项关键挑战,现有方法常因过度防御而损害模型效用,或依赖浅层对齐,难以识别需深度推理的复杂威胁。为此,我们提出PRISM(多模态集成安全的原理性推理),一个类系统2的框架,通过结构化的四阶段推理过程,明确应对三类不同的多模态安全违规。该框架包含两个核心组件:专用阶段分析各类违规的结构化推理流水线,以及通过蒙特卡洛树搜索(MCTS)生成的PRISM-DPO,利用直接偏好优化(DPO)提升推理质量。全面评估表明,PRISM显著降低了在JailbreakV-28K和VLBreak上的攻击成功率,增强了对抗自适应攻击的鲁棒性,并能泛化至分布外的多图像威胁,同时在良性多模态基准上更好保留模型效用。代码、数据及模型权重已公开于https://github.com/SaFoLab-WISC/PRISM。
原文摘要 · Abstract (English)
Safeguarding vision-language models (VLMs) is a critical challenge, as existing methods often suffer from over-defense, which harms utility, or rely on shallow alignment, failing to detect complex threats that require deep reasoning. To this end, we introduc PRISM (Principled Reasoning for Integrated Safety in Multimodality), a System 2-like framework that aligns VLMs through a structured four-stage reasoning process explicitly designed to handle three distinct categories of multimodal safety violations. Our framework consists of two key components: a structured reasoning pipeline that analyzes each violation category in dedicated stages, and PRISM-DPO, generated via Monte Carlo Tree Search (MCTS) to refine reasoning quality through Direct Preference Optimization. Comprehensive evaluations show that PRISM substantially reduces attack success rates on JailbreakV-28K and VLBreak, improves robustness against adaptive attacks, and generalizes to out-of-distribution multi-image threats, while better preserving model utility on benign multimodal benchmarks. Our code, data, and model weights available at https://github.com/SaFoLab-WISC/PRISM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。