用图像提示动态调整模型安全,减少误拒并适配不同价值观。
Reimagining Safety Alignment with An Image
- 通过优化图像提示,让模型在不更新参数下适应不同安全偏好。
- 在多数据集上实现更优的安全性与可用性平衡,误拒率显著降低。
- 适合需要灵活安全对齐的多模态模型部署场景。
大型语言模型(LLMs)在各类应用中表现优异,但面临双重挑战:在越狱攻击下生成有害内容,以及因僵化安全机制而过度拒绝正常请求。这些问题在多模态大模型(MLLMs)中尤为突出,表现为跨模态任务中过量拒绝及攻击面扩大带来的新安全风险。传统方法如SFT和RLHF因需昂贵的参数调优且无法支持单一模型中的多重价值体系,难以应对。本文提出Magic Image——一种基于优化的视觉提示框架,通过使用有害/良性样本优化图像提示,使单个模型能够适应不同价值体系,并在不修改参数的情况下更好对齐给定的安全偏好。实验表明,该方法在多个数据集上提升了安全与有效性的平衡,同时保持模型性能,为可部署的多模态模型安全对齐提供了实用解决方案。
原文摘要 · Abstract (English)
Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusal of benign queries due to rigid safety mechanisms. These issues are further complicated by the need to accommodate different value systems and precisely align with given safety preferences. Moreover, traditional methods like SFT and RLHF lack this capability due to their costly parameter tuning requirements and inability to support multiple value systems within a single model. These problems are more obvious in multimodal large language models (MLLMs), especially in terms of heightened over-refusal in cross-modal tasks and new security risks arising from expanded attack surfaces. We propose Magic Image, an optimization-driven visual prompt framework that enhances security while reducing over-refusal. By optimizing image prompts using harmful/benign samples, our method enables a single model to adapt to different value systems and better align with given safety preferences without parameter updates. Experiments demonstrate improved safety-effectiveness balance across diverse datasets while preserving model performance, offering a practical solution for deployable MLLM safety alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。