分析CLIP模型中社会偏见如何从全局到局部传递,揭示其迁移规律。
From Global to Local: Social Bias Transfer in CLIP
- 对比全局与局部数据视角下的预训练偏见差异。
- 发现下游任务偏见与预训练偏见无稳定相关性。
- 解释偏见不一致传递的原因:适配后表示空间趋于收敛。
将对比语言-图像预训练(CLIP)模型作为大量下游任务的主干网络,引发了对其可迁移性影响的深入分析,尤其是其众所周知的社会偏见和人类刻板印象的再现。这些在预训练阶段学习到的偏见,会如何传播至视觉问答或图像字幕等下游应用?它们是否真的会发生转移?我们通过全面的实证分析探讨了这一现象,即先前文献中提到的偏见迁移。首先,我们考察了数据全局与局部视图下预训练偏见的变化,发现偏见度量高度依赖于计算所用的数据子集。其次,我们分析了预训练模型偏见与下游任务偏见在不同预训练偏见水平下的相关性,发现难以发现一致的偏见迁移趋势。最后,我们探究了这种不一致性产生的原因,表明在当前范式下,不同预训练的CLIP模型在适配下游任务时,其表示空间趋向收敛。我们希望本研究能为偏见行为提供洞察,并推动未来更优的偏见缓解实践。
原文摘要 · Abstract (English)
The recycling of contrastive language-image pre-trained (CLIP) models as backbones for a large number of downstream tasks calls for a thorough analysis of their transferability implications, especially their well-documented reproduction of social biases and human stereotypes. How do such biases, learned during pre-training, propagate to downstream applications like visual question answering or image captioning? Do they transfer at all? We investigate this phenomenon, referred to as bias transfer in prior literature, through a comprehensive empirical analysis. Firstly, we examine how pre-training bias varies between global and local views of data, finding that bias measurement is highly dependent on the subset of data on which it is computed. Secondly, we analyze correlations between biases in the pre-trained models and the downstream tasks across varying levels of pre-training bias, finding difficulty in discovering consistent trends in bias transfer. Finally, we explore why this inconsistency occurs, showing that under the current paradigm, representation spaces of different pre-trained CLIPs tend to converge when adapted for downstream tasks. We hope this work offers valuable insights into bias behavior and informs future research to promote better bias mitigation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。