arXiv:2502.07957cs.AI2025-02NAACL被引 7

CLIP模型的内在偏见主要由预训练数据决定,且与下游性能强相关。

Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders

  • 分析131个CLIP模型,发现预训练数据是影响偏见的最关键因素。
  • 偏见程度与下游性能相关性达0.3至0.8,性能越高偏见越明显。
  • 跨模态测试显示社会群体偏见依赖于视觉或语言模态,需分模态治理。

尽管近期研究发现基于对比语言图像预训练(CLIP)框架的视觉-语言模型存在内在社会偏见,但其上游预训练特征如何影响偏见,以及偏见与下游性能的关系仍不清晰。本文对131个在26个数据集上、使用55种架构、多种规模的CLIP模型进行了迄今最全面的分析,通过26项成熟的单模态与跨模态嵌入关联测试评估每种模型的偏见。结果表明,预训练数据的选择是预测偏见的最主要上游因素,而模型架构差异影响甚微。此外,采用复杂过滤技术以提升下游性能的数据集往往伴随更高水平的内在偏见。我们还发现,内在偏见与下游性能显著相关(0.3 ≤ r ≤ 0.8),说明为性能优化的模型无意中放大了表征偏见。单模态与跨模态测试的比较揭示,社会群体偏见高度依赖模态类型。研究提示:必须在模型全开发周期内采取更精细的策略应对视觉-语言模型的内在偏见。

原文摘要 · Abstract (English)

While recent work has found that vision-language models trained under the Contrastive Language Image Pre-training (CLIP) framework contain intrinsic social biases, the extent to which different upstream pre-training features of the framework relate to these biases, and hence how intrinsic bias and downstream performance are connected has been unclear. In this work, we present the largest comprehensive analysis to-date of how the upstream pre-training factors and downstream performance of CLIP models relate to their intrinsic biases. Studying 131 unique CLIP models, trained on 26 datasets, using 55 architectures, and in a variety of sizes, we evaluate bias in each model using 26 well-established unimodal and cross-modal principled Embedding Association Tests. We find that the choice of pre-training dataset is the most significant upstream predictor of bias, whereas architectural variations have minimal impact. Additionally, datasets curated using sophisticated filtering techniques aimed at enhancing downstream model performance tend to be associated with higher levels of intrinsic bias. Finally, we observe that intrinsic bias is often significantly correlated with downstream performance ($0.3 \leq r \leq 0.8$), suggesting that models optimized for performance inadvertently learn to amplify representational biases. Comparisons between unimodal and cross-modal association tests reveal that social group bias depends heavily on the modality. Our findings imply that more sophisticated strategies are needed to address intrinsic model bias for vision-language models across the entire model development pipeline.

CLIP偏见分析多模态模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。