arXiv:2509.07596cs.CV2025-09ICCV被引 6

发现视觉语言模型性别偏见评估受无关特征干扰,结果不可靠。

Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation

  • 通过扰动物体和背景等非性别特征,检验其对偏见评估的影响。
  • 仅10%物体遮挡或弱模糊背景,偏见分数最高变化175%。
  • 建议同时报告偏见得分与特征敏感度,提升评估可信度。

视觉语言基础模型(VLMs)中的性别偏见引发对其安全部署的担忧,通常通过带有性别标注的真实图像基准进行评估。然而,这些基准常存在性别与非性别特征(如物体、背景)之间的虚假相关性。本文系统地在四个常用基准(COCO-gender、FACET、MIAP、PHASE)及多种VLMs上扰动非性别特征,量化其对偏见评估的影响。结果显示,即使仅遮挡10%物体或弱模糊背景,偏见分数仍可发生剧烈变化:生成型VLMs最高波动达175%,CLIP变体达43%。这表明当前评估结果往往反映模型对虚假特征的响应,而非真实性别偏见,严重削弱评估可靠性。由于完全消除虚假特征的基准难以构建,本文建议在报告偏见指标的同时,附加特征敏感度测量,以实现更可靠的偏见评估。

原文摘要 · Abstract (English)

Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such as objects and backgrounds, we identify a critical oversight in gender bias evaluation: Do spurious features distort gender bias evaluation? To address this question, we systematically perturb non-gender features across four widely used benchmarks (COCO-gender, FACET, MIAP, and PHASE) and various VLMs to quantify their impact on bias evaluation. Our findings reveal that even minimal perturbations, such as masking just 10% of objects or weakly blurring backgrounds, can dramatically alter bias scores, shifting metrics by up to 175% in generative VLMs and 43% in CLIP variants. This suggests that current bias evaluations often reflect model responses to spurious features rather than gender bias, undermining their reliability. Since creating spurious feature-free benchmarks is fundamentally challenging, we recommend reporting bias metrics alongside feature-sensitivity measurements to enable a more reliable bias assessment.

性别偏见基准评估虚假相关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。