arXiv:2605.00326cs.CLcs.CV2026-05

同一图像用不同等效提示,安全得分差异大,提示变动能暴露模型不靠谱。

Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification

论文配图:Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification
图 1 · 摘自论文原文
  • 用多个等效提示平均得分,提升零样本安全分类可靠性。
  • 跨提示得分方差越大,错误率越高,可作模型脆弱性诊断。
  • 无需训练的平均集成法优于调参校准,适合无标签场景。

零样本视觉语言模型(VLM)的安全分类常依赖单个提示的第一个词概率作为决策分数,但我们发现:即使二分类输出位置固定,语义等价的提示改写也会导致同一样本的不安全概率显著变化。在多组多模态安全基准和多种VLM家族中,跨提示得分方差与提示间分歧及更高错误率强相关,可作为模型脆弱性的有效诊断指标。一种无需训练的均值集成方法,在全部14个数据集-模型组合上降低NLL,12/14组合降低ECE,头对头比较胜过温度校准、Platt校准和等距回归。在AUROC和AUPRC上,相对于选定最优提示基线均有排名提升;与完整15提示分布相比,AUPRC仍保持一致,而AUROC略有下降。有标签时,在均值基础上再做标注校准可进一步提升性能。我们将其视为对零样本首词安全分数可靠性的压力测试,建议以提示族平均作为无标签场景下的标准可靠性基线。

原文摘要 · Abstract (English)

Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they are unreliable under semantically equivalent prompt reformulation: even when the binary label is constrained to a fixed output position, equivalent prompts can induce materially different unsafe probabilities for the same sample. Across multimodal safety benchmarks and multiple VLM families, cross-prompt variance is strongly associated with prompt-level disagreement and higher error, making it a useful fragility diagnostic. A training-free mean ensemble improves NLL on all 14 dataset-model evaluation pairs and ECE on 12/14 relative to a train-selected single-prompt baseline, and wins more head-to-head NLL comparisons than labeled temperature scaling, Platt scaling, and isotonic regression applied to the same prompt. Ranking gains are consistent against the train-selected baseline on both AUROC and AUPRC, and against the full 15-prompt distribution remain consistent on AUPRC while softening on AUROC. Labeled calibration on top of the mean provides further gains when labels are available, identifying prompt averaging as a strong label-free first stage rather than a replacement for calibration. We frame this as a reliability stress test for zero-shot VLM first-token safety scores and recommend prompt-family evaluation with mean aggregation as a standard label-free reliability baseline.

零样本安全分类提示工程可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。