提出盲偏好优化方法,提升多模态模型安全性和抗风险能力
Towards Harmless Multimodal Assistants with Blind Preference Optimization
- 通过构建多模态安全偏好数据集,设计盲偏好优化策略
- 使基础模型安全率提升45.0%,在多个评测集上显著降低不安全响应
- 适合关注多模态模型安全、对齐与防御的研究者使用
多模态大语言模型(MLLMs)在多模态理解、推理和交互方面展现出强大能力。随着其广泛应用,安全性问题日益突出。鉴于偏好优化在对齐人类偏好方面的有效性,亟需面向安全的多模态偏好数据。为此,我们构建了MMSafe-PO偏好数据集,包含多模态指令、对话格式及基于人类反馈的成对响应排序。我们发现两个关键现象:模态协同防御与模态欺骗,表明MLLM具备一定内在防御能力,但仍存在独特安全挑战。基于此,提出盲偏好优化(BPO)方法。在三个基准上的实验表明,BPO有效提升了MLLM的安全性。显著将基线模型安全率提高45.0%,优于DPO方法。将BPO应用于MMSafe-PO数据集,可使基线模型在其他安全评测集上的不安全率分别降至14.5%(MM-SafetyBench)和82.9%(HarmEval),验证了数据集与方法的有效性与鲁棒性。代码与数据已开源。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLMs, the associated safety issues have become increasingly critical. Due to the effectiveness of preference optimization in aligning MLLMs with human preferences, there is an urgent need for safety-related preference data for MLLMs. To address this, we construct the MMSafe-PO preference dataset towards harmless multimodal assistants, featuring multimodal instructions, the conversational format, and ranked paired responses from human feedback. We also identify two insightful observations: modality co-defense and modality cheating, which illustrate that MLLMs possess a certain level of inherent defense while still presenting unique safety challenges. Based on these observations, we propose the Blind Preference Optimization (BPO) approach. Comprehensive experiments on three benchmarks show that BPO effectively enhances the safety capabilities of MLLMs. Notably, BPO significantly improves the safety rate of the base MLLM by 45.0%, outperforming the DPO approach. Additionally, applying BPO to the MMSafe-PO dataset greatly reduces the base MLLM's unsafe rate on other safety benchmarks (14.5% on MM-SafetyBench and 82.9% on HarmEval, demonstrating the effectiveness and robustness of both the dataset and the approach. We release code and data at https://lu-yang666.github.io/MMsafe-PO-Web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。