用图像描述强化学习提升视觉语言模型安全性
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

- 通过自生成图像描述引导模型安全决策
- 在五个安全基准上提升3.7至19.0分
- 适合研究多模态安全对齐的学者
大型视觉语言模型(LVLMs)仍易受越狱攻击,此类攻击利用视觉输入绕过语言模型原有的安全对齐。我们提出SafeCap,一种基于强化学习的框架,通过学习到的自描述机制对齐LVLM。SafeCap训练一个策略模型先生成与安全相关的图像描述,再生成最终回答;该描述通过冻结的LLM是否能做出安全决策来优化。这种以描述为中介的目标促使策略暴露有助于安全响应生成的视觉线索,而非仅依赖直接拒绝监督。在五个多模态安全基准和六个视觉效用基准上,SafeCap在设定的DirectCap协议下显著提升整体安全性能,在四种模型设置中安全平均分提升3.7至19.0点,同时保持相当或更高的视觉效用。在匹配骨干和数据的对照实验中,SafeCap优于安全SFT、DPO和SafeGRPO,证明了描述中介强化学习在多模态安全对齐中的有效性。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。