arXiv:2411.00997cs.CVcs.AI2024-11AAAI被引 52

发现CLIP模型对特定群体存在隐性社会偏见,如将中东男性与'恐怖分子'关联。

Identifying Implicit Social Biases in Vision-Language Models

  • 构建包含374个词的偏见分类体系So-B-IT,分析图文交互中的社会偏见
  • 在提示中使用

视觉语言模型(如CLIP)在多模态检索任务中日益流行。然而,已有研究显示,大型语言和深度视觉模型可能习得训练数据中的历史偏见,导致刻板印象延续并造成潜在下游危害。本文系统分析了CLIP中存在的社会偏见,重点关注图像与文本模态的交互。我们提出一种名为So-B-IT的偏见分类体系,涵盖374个词语,分为十类偏见类型,每类若与特定人口群体关联,均可能导致社会伤害。利用该分类体系,我们以每个词作为提示词,从人脸图像数据集中检索CLIP生成的图像。结果发现,当提示为“恐怖分子”时,CLIP频繁检索出中东男性图像,表现出明显的不良关联。进一步分析表明,这些有害刻板印象同样存在于用于训练CLIP的大规模图像-文本数据集中。研究强调评估与缓解视觉语言模型偏见的重要性,并呼吁对大规模预训练数据集进行透明化、公平性导向的筛选与管理。

原文摘要 · Abstract (English)

Vision-language models, like CLIP (Contrastive Language Image Pretraining), are becoming increasingly popular for a wide range of multimodal retrieval tasks. However, prior work has shown that large language and deep vision models can learn historical biases contained in their training sets, leading to perpetuation of stereotypes and potential downstream harm. In this work, we conduct a systematic analysis of the social biases that are present in CLIP, with a focus on the interaction between image and text modalities. We first propose a taxonomy of social biases called So-B-IT, which contains 374 words categorized across ten types of bias. Each type can lead to societal harm if associated with a particular demographic group. Using this taxonomy, we examine images retrieved by CLIP from a facial image dataset using each word as part of a prompt. We find that CLIP frequently displays undesirable associations between harmful words and specific demographic groups, such as retrieving mostly pictures of Middle Eastern men when asked to retrieve images of a "terrorist". Finally, we conduct an analysis of the source of such biases, by showing that the same harmful stereotypes are also present in a large image-text dataset used to train CLIP models for examples of biases that we find. Our findings highlight the importance of evaluating and addressing bias in vision-language models, and suggest the need for transparency and fairness-aware curation of large pre-training datasets.

视觉语言模型社会偏见公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。