arXiv:2608.23313cs.AI2026-08

让视觉语言模型的安全性评估从‘是否拒绝’升级到‘为何拒绝’。

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

论文配图:EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models
图 1 · 摘自论文原文
  • 通过多轮提问检测模型是否基于图文证据做出安全判断。
  • 实测11个模型在正确识别危险上的准确率仅27.6%~52.8%。
  • 适合关注模型真实安全机制的研究者与开发者。

视觉语言模型的安全评估通常只看最终回应:是否拒绝、警告或遵从。这种结果层面的评价无法判断模型是否因正确的多模态原因而安全。看似安全的行为可能源于关键词触发拒绝、遗漏视觉风险,或对无害敏感输入过度拒绝。我们提出EviSafe,一种基于证据的VLM安全评估框架,联合评估用户面对行为、显式图文证据依赖性,以及对关键证据反事实变化的敏感性。EviSafeBench将该框架实现为受控基准,包含1,181个黄金图文场景和2,452个针对性反事实变体,覆盖八个安全领域和八种风险源类型。每个场景包含黄金安全决策、证据标注、安全响应策略及反事实干预。三阶段探针协议通过自然回应、证据报告和反事实回应提示,使用基于证据的评判器打分。在11个评估的VLM中,自然严重性准确率介于27.6%至52.8%,宽松诊断一致性为6.1%至29.3%,从不安全到安全的反事实转变成功率在30.4%至58.4%之间。这些差距表明,当前模型并未可靠地因正确多模态原因而安全,推动评估超越单纯拒绝次数。

原文摘要 · Abstract (English)

Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view cannot tell whether a model is safe for the right multimodal reason. Safelooking behavior may reflect keyword-triggered refusal, missed visual hazards, or over-refusal of benign-sensitive inputs. We introduce EviSafe, an evidence-grounded framework for VLM safety that jointly evaluates natural user-facing behavior, explicit grounding in textual and visual evidence, and behavioral sensitivity to counterfactual changes in safety-critical evidence. EviSafeBench instantiates the framework as a controlled benchmark with 1,181 gold image-text scenarios and 2,452 targeted counterfactual variants across eight safety domains and eight risk-source types. Each scenario includes a gold safety decision, evidence annotations, a safe-response policy, and counterfactual interventions. The three-probe protocol queries models with natural-response, evidencereporting, and counterfactual-response prompts, then scores them using an evidence-aware judge. Across eleven evaluated VLMs, natural severity accuracy ranges from 27.6% to 52.8%, relaxed diagnostic consistency from 6.1% to 29.3%, and unsafe-to-safe counterfactual transition success from 30.4% to 58.4%. These gaps show that the evaluated VLMs are not reliably safe for the right multimodal reason and motivate evaluation beyond refusal counts.

多模态安全评估框架视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。