arXiv:2503.06223cs.CV2025-03

用扩散模型生成有害视觉上下文,发现大模型在多模态下安全失效的隐藏漏洞。

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

  • 通过强化学习驱动扩散模型生成带毒文本的视觉上下文,实现黑盒安全测试。
  • 在LLaVA上使不安全响应率提升最高达10.69%,且对Gemini等模型有强迁移性。
  • 揭示当前安全防护机制对真实多模态风险无效,适合关注模型安全的研究者。

大型视觉语言模型(VLMs)日益部署于开放环境,确保多模态输入下的可靠安全性至关重要。然而现有评估仍以指令为中心,聚焦明确恶意查询,忽视了更现实且未被充分探索的风险:模型在有害上下文暴露下安全性对齐是否依然稳健。这一局限对多模态系统尤为关键,因视觉输入可显著引导模型行为,使仅依赖文本的审计不足。本文研究在有害上下文暴露下的多模态安全审计,探讨当部分有毒文本与视觉上下文结合时,VLMs能否保持安全行为。为此,我们提出RedDiff (RedDiffuser),一种基于强化学习的框架,利用扩散模型生成语义连贯的视觉输入,用于黑盒安全测试。通过贪心提示搜索与强化优化结合,RedDiffuser揭示了高风险多模态输入,暴露出潜在安全缺陷。在开源与商用VLM上的大量实验表明,此类上下文相关失败普遍存在。在LLaVA上,RedDiffuser使原始集不安全响应率提升最高达10.69%,持留集提升8.91%,且对Gemini和LLaMA-Vision具有强迁移性。这些漏洞即使在外部安全防护下仍存在,表明当前系统级安全机制对真实多模态风险仍显不足。研究揭示了现有安全评估的关键盲点,并确立上下文感知的多模态审计作为诊断现代VLM系统隐藏漏洞的必要范式。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is critical. However, existing evaluations remain largely instruction-centric, focusing on explicit malicious queries while overlooking a more realistic and underexplored risk: whether safety alignment remains robust under harmful contextual exposure. This limitation is particularly important for multimodal systems, where visual inputs can substantially steer model behavior and render text-only auditing insufficient. In this work, we study multimodal safety auditing under harmful contextual exposure, asking whether VLMs can maintain safe behavior when partial toxic text is paired with visual context. To enable systematic auditing, we propose RedDiffuser (RedDiff), a reinforcement-based framework that leverages diffusion models to generate semantically coherent visual inputs for black-box safety testing. By combining greedy prompt search with reinforcement optimization, RedDiffuser uncovers high-risk multimodal inputs that expose latent safety failures. Extensive experiments on both open-source and commercial VLMs show that such context-conditioned failures are widespread. On LLaVA, RedDiffuser increases unsafe response rates by up to 10.69% on the original set and 8.91% on a hold-out set, with strong transferability to Gemini and LLaMA-Vision. These vulnerabilities persist even under external safety guardrails, suggesting that current system-level safety mechanisms remain insufficient for realistic multimodal risks. Our findings reveal a critical blind spot in existing safety evaluations and establish context-aware multimodal auditing as an essential paradigm for diagnosing hidden vulnerabilities in modern VLM systems.

多模态安全扩散模型模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。