arXiv:2412.18826cs.CL2024-12被引 22

用动态提示提升多模态模型安全,防止生成有害内容。

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting

  • 基于多模态思维链生成场景化安全提示
  • 在多个基准上显著减少有害输出且不影响正常任务表现
  • 适合关注多模态模型安全的开发者与研究者

尽管多模态大语言模型(MLLMs)在视觉-语言推理方面取得显著进展,但相比仅处理文本的模型,其生成有害内容的风险更高。现有防御性提示技术依赖静态统一的安全准则,无法适应不同多模态场景的特异性风险。为此,我们提出RapGuard,一种利用多模态思维链推理动态生成场景化安全提示的新框架。RapGuard通过针对每个输入的独特风险自适应调整提示,有效缓解有害输出,同时保持良性任务的高性能。在多个MLLM基准上的实验结果表明,RapGuard实现了领先的安全部分表现,显著降低有害内容产出,且不损害响应质量。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have made remarkable progress in vision-language reasoning, they are also more susceptible to producing harmful content compared to models that focus solely on text. Existing defensive prompting techniques rely on a static, unified safety guideline that fails to account for the specific risks inherent in different multimodal contexts. To address these limitations, we propose RapGuard, a novel framework that uses multimodal chain-of-thought reasoning to dynamically generate scenario-specific safety prompts. RapGuard enhances safety by adapting its prompts to the unique risks of each input, effectively mitigating harmful outputs while maintaining high performance on benign tasks. Our experimental results across multiple MLLM benchmarks demonstrate that RapGuard achieves state-of-the-art safety performance, significantly reducing harmful content without degrading the quality of responses.

多模态安全防御提示链式思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。