arXiv:2503.07878cs.CVcs.AI2025-03中稿 · WACV 2026被引 1

提出新指标DBAC,精准识别图像描述中偏见放大的源头。

A Woman with a Knife or A Knife with a Woman? Measuring Directional Bias Amplification in Image Captions

  • 设计语言感知的定向偏见度量方法,定位偏见放大位置。
  • 在COCO数据集上验证,仅DBAC能可靠测量描述中的偏见放大。
  • 对句子编码器不敏感,比现有方法更稳定准确,适合研究视觉-语言模型偏见。

当模型在有偏见的数据集上训练时,不仅会复现数据中的偏见,还可能在测试时加剧这种偏见,这一现象称为偏见放大。当前多数偏见放大度量(如BA(MALS)、DPA)仅适用于分类数据集,无法捕捉图像描述中的语言语义。近期工作提出了语言感知的偏见放大度量LIC,可理解描述语义,但存在关键缺陷:无法定位偏见放大的来源。本文提出图像描述中的方向性偏见放大度量(DBAC),是一种语言感知且具备方向性的新指标,能够识别描述模型何时放大偏见。相较于LIC,DBAC具有两项改进:(1) 对句子编码器的敏感度更低(语言感知度量中的超参数);(2) 对描述中偏见放大的估计更准确。在包含性别与种族属性的COCO图像描述数据集上的实验表明,只有DBAC能可靠测量描述中的偏见放大。

原文摘要 · Abstract (English)

When we train models on biased datasets, they not only reproduce data biases, but can worsen them at test time - a phenomenon called bias amplification. Many of the current bias amplification metrics (e.g., BA (MALS), DPA) measure bias amplification only in classification datasets. These metrics are ineffective for image captioning datasets, as they cannot capture the language semantics of a caption. Recent work introduced Leakage in Captioning (LIC), a language-aware bias amplification metric that understands caption semantics. However, LIC has a crucial limitation: it cannot identify the source of bias amplification in captioning models. We propose Directional Bias Amplification in Captioning (DBAC), a language-aware and directional metric that can identify when captioning models amplify biases. DBAC has two more improvements over LIC: (1) it is less sensitive to sentence encoders (a hyperparameter in language-aware metrics), and (2) it provides a more accurate estimate of bias amplification in captions. Our experiments on gender and race attributes in the COCO captions dataset show that DBAC is the only reliable metric to measure bias amplification in captions.

偏见放大图像描述语言感知评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。