评测视觉语言模型在双用途图像下的过度拒绝与安全完成能力
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
- 构建多模态基准,评估模型在含危险图像的良性请求下的表现
- 最佳模型仅12.9%实现安全完成,多数模型要么过度拒绝要么放任风险
- 适合关注AI安全对齐、多模态内容审核的研究者与开发者
随着视觉语言模型(VLMs)能力不断提升,如何在安全与实用性间保持平衡成为核心挑战。安全机制虽必要,却可能引发过度拒绝——即对无害请求也予以拒绝。然而,当前缺乏系统性评估视觉模态下过度拒绝的基准。此类场景具有独特挑战,例如指令本身无害但图像含危险内容(双用途情况)。模型在此类情形中常表现不佳:或过度保守拒绝,或不加甄别生成危险内容,凸显需更精细的安全对齐。理想行为应为安全完成——执行无害部分的同时明确警告潜在危害。为此,我们提出DUAL-Bench,一个大规模多模态基准,聚焦于评估VLMs的过度拒绝与安全完成能力。我们在12种危险类别下,对18个VLMs进行评估,采用语义保持的视觉扰动。结果显示,模型在双用途场景中安全性边界极脆弱,陷入二元困境:非过度敏感拒绝即完全失控生成。即使最优模型GPT-5-Nano,安全完成率也仅为12.9%,而GPT-5和Qwen系列平均仅7.9%与3.9%。我们希望DUAL-Bench推动更精细的多模态安全对齐策略。
原文摘要 · Abstract (English)
As vision-language models (VLMs) become increasingly capable, maintaining a balance between safety and usefulness remains a central challenge. Safety mechanisms, while essential, can backfire, causing over-refusal, where models decline benign requests out of excessive caution. Yet, there is currently a significant lack of benchmarks that have systematically addressed over-refusal in the visual modality. This setting introduces unique challenges, such as dual-use cases where an instruction is harmless, but the accompanying image contains harmful content. Models frequently fail in such scenarios, either refusing too conservatively or completing tasks unsafely, which highlights the need for more fine-grained alignment. The ideal behaviour is safe completion, i.e., fulfilling the benign parts of a request while explicitly warning about any potentially harmful elements. To address this, we present DUAL-Bench, a large scale multimodal benchmark focused on over-refusal and safe completion in VLMs. We evaluated 18 VLMs across 12 hazard categories under semantics-preserving visual perturbations. In dual-use scenarios, models exhibit extremely fragile safety boundaries. They fall into a binary trap: either overly sensitive direct refusal or defenseless generation of dangerous content. Consequently, even the best-performing model GPT-5-Nano, at just 12.9% safe completion, with GPT-5 and Qwen families averaging 7.9% and 3.9%. We hope DUAL-Bench fosters nuanced alignment strategies balancing multimodal safety and utility. Content Warning: This paper contains examples of sensitive and potentially hazardous content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。