arXiv:2510.13190cs.CL2025-10被引 5

用细粒度分类和定向引导提升视觉语言模型的安全性

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

  • 通过分类与定向提示结合,动态生成安全响应
  • 在5个基准上降低越狱率与不遵从率
  • 无需重训练,适合各类对齐程度的模型

大型视觉语言模型(LVLMs)虽具备强大多模态推理能力,但也扩大了攻击面,尤其体现在恶意目标隐藏于看似无害的提示中。我们提出SHIELD,一种轻量级、模型无关的预处理框架,将细粒度安全分类与类别特定引导及明确动作(阻断、重构、放行)相结合。相比二元审核器,SHIELD能生成定制化安全提示,实现细腻拒绝或安全引导,无需重新训练。在五个基准和五种代表性LVLM上,该方法持续降低越狱率与不遵从率,同时保持模型可用性。本方法即插即用,开销极小,可轻松扩展至新攻击类型,为弱对齐与强对齐的LVLM提供实用安全补丁。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) unlock powerful multimodal reasoning but also expand the attack surface, particularly through adversarial inputs that conceal harmful goals in benign prompts. We propose SHIELD, a lightweight, model-agnostic preprocessing framework that couples fine-grained safety classification with category-specific guidance and explicit actions (Block, Reframe, Forward). Unlike binary moderators, SHIELD composes tailored safety prompts that enforce nuanced refusals or safe redirection without retraining. Across five benchmarks and five representative LVLMs, SHIELD consistently lowers jailbreak and non-following rates while preserving utility. Our method is plug-and-play, incurs negligible overhead, and is easily extendable to new attack types -- serving as a practical safety patch for both weakly and strongly aligned LVLMs.

视觉语言模型安全防御提示工程鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。