不用复杂恶意数据,用简单拒绝语就能提升多模态模型安全
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
- 用良性指令数据替换响应为明确拒绝语进行微调
- 只需少量拒绝数据即可显著提升模型安全性
- 适合关注多模态安全但缺乏标注资源的研究者
多模态大语言模型(MLLM)虽取得进展,但其安全对齐仍受限。现有开源模型依赖语言模块的对齐能力,却缺乏针对多模态输入的安全机制,易受视觉攻击(如字形篡改)。当前方法依赖精心设计的安全数据集,但其有效知识尚不明确。对比实验表明,安全差距主要源于数据分布偏差,而非图像内容、回复质量或数据集对比行为。为此,我们提出在小规模良性指令数据上微调,将响应替换为简洁的拒绝语。实验显示,无需人工构建高质量恶意数据,仅需在微调集中包含特定比例的拒绝数据,即可显著提升模型安全性,说明安全对齐未丢失,而是被多模态预训练或指令微调掩盖。纠正底层数据偏差即可缩小视觉域安全差距。
原文摘要 · Abstract (English)
Multi-modal large language models (MLLMs) have made significant progress, yet their safety alignment remains limited. Typically, current open-source MLLMs rely on the alignment inherited from their language module to avoid harmful generations. However, the lack of safety measures specifically designed for multi-modal inputs creates an alignment gap, leaving MLLMs vulnerable to vision-domain attacks such as typographic manipulation. Current methods utilize a carefully designed safety dataset to enhance model defense capability, while the specific knowledge or patterns acquired from the high-quality dataset remain unclear. Through comparison experiments, we find that the alignment gap primarily arises from data distribution biases, while image content, response quality, or the contrastive behavior of the dataset makes little contribution to boosting multi-modal safety. To further investigate this and identify the key factors in improving MLLM safety, we propose finetuning MLLMs on a small set of benign instruct-following data with responses replaced by simple, clear rejection sentences. Experiments show that, without the need for labor-intensive collection of high-quality malicious data, model safety can still be significantly improved, as long as a specific fraction of rejection data exists in the finetuning set, indicating the security alignment is not lost but rather obscured during multi-modal pretraining or instruction finetuning. Simply correcting the underlying data bias could narrow the safety gap in the vision domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。