arXiv:2411.01703cs.CLcs.AI2024-11被引 21

UniGuard为多模态大模型提供通用安全防护,抵御越狱攻击。

UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models

  • 联合分析单模态与跨模态有害信号,构建统一防御机制。
  • 在多个大模型和攻击策略下有效降低有害响应概率。
  • 部署轻量,不影响模型原有的视觉语言理解能力。

多模态大语言模型(MLLMs)虽革新了视觉-语言理解,但仍易受多模态越狱攻击影响,攻击者通过精心设计的输入诱导模型生成有害或不当内容。本文提出UniGuard,一种新型多模态安全防护机制,同时考虑单模态与跨模态有害信号。UniGuard通过在有毒语料上训练多模态防护器,以最小化生成有害回应的概率。该防护器可在推理阶段无缝应用于任意输入提示,计算开销极低。大量实验表明,UniGuard在多种模态、攻击策略及多个前沿MLLM(包括LLaVA、Gemini Pro、GPT-4o、MiniGPT-4、InstructBLIP)上均表现出良好泛化性,且保持模型整体视觉-语言理解能力。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have revolutionized vision-language understanding but remain vulnerable to multimodal jailbreak attacks, where adversarial inputs are meticulously crafted to elicit harmful or inappropriate responses. We propose UniGuard, a novel multimodal safety guardrail that jointly considers the unimodal and cross-modal harmful signals. UniGuard trains a multimodal guardrail to minimize the likelihood of generating harmful responses in a toxic corpus. The guardrail can be seamlessly applied to any input prompt during inference with minimal computational costs. Extensive experiments demonstrate the generalizability of UniGuard across multiple modalities, attack strategies, and multiple state-of-the-art MLLMs, including LLaVA, Gemini Pro, GPT-4o, MiniGPT-4, and InstructBLIP. Notably, this robust defense mechanism maintains the models' overall vision-language understanding capabilities.

多模态安全防护越狱攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。