分析LAION-5B数据集中的年龄、性别、种族与情绪偏见,揭示其对AI系统的影响。
Unmasking LAION-5B: Age, Gender, Race, and Emotion Biases in Large-Scale Image Datasets

- 用FairFace、DeepFace等模型分析图像中人脸的年龄性别种族与情绪
- 年轻成年白人男性显著过量,中老年女性及少数族裔严重不足
- 发现性别与情绪的刻板关联,适合关注AI公平性的研究者阅读
大规模图像文本数据集如LAION-5B是现代AI系统的基础,但其海量且未经筛选的特性引发了对人口统计与刻板印象偏见的重大担忧。本研究对LAION-5B的两大组成部分LAION-2B-en和LAION-2B-multi进行了全面分析,使用FairFace、DeepFace和Emo-AffectNet等先进模型检测图像中的人脸,识别年龄、性别、种族和情绪表达方面的偏见。结果表明,年轻成年人(20-39岁)、白人和男性在两个数据集中均显著过量,而少数族裔及中老年女性则持续被低估。此外,还观察到与性别相关的刻板情绪关联,例如‘愤怒’多与男性相关,‘快乐’多与女性相关,揭示了数据中的系统性不平衡。这些模式在两种人口属性模型和两个数据集组件中高度一致,说明这些偏见已深度嵌入当前最广泛使用的训练数据集中。鉴于LAION-5B被大规模用于生成模型训练,这些人口统计偏差可能影响众多下游AI系统的性能与行为。
原文摘要 · Abstract (English)
Large-scale image-text datasets, such as LAION-5B, are foundational to modern AI systems, yet their vast scale and uncurated nature raise significant concerns about demographic and stereotypical biases. This study presents a comprehensive analysis of the demographic composition and representational, stereotypical, and intersectional biases in LAION-2B-en and LAION-2B-multi, the two main components of the LAION-5B dataset. Using state-of-the-art models -- FairFace, DeepFace, and Emo-AffectNet -- we analyze faces detected in the dataset to identify biases across age, gender, race, and expressed emotion. Our findings reveal substantial overrepresentation of young adults (20--39), White individuals, and males, alongside consistent underrepresentation of minority racial groups and middle-aged or older women across both dataset components. We also observe stereotypical associations between demographic attributes and emotions, such as ``Anger'' being predominantly linked to males and ``Happiness'' to females, pointing to systemic imbalances in the data. The consistency of these patterns across two demographic models and both components of LAION-5B demonstrates that these biases are deeply embedded in one of the most widely-used training datasets. Given the scale at which LAION-5B is used to train generative models, these demographic imbalances could shape the behavior and outputs of numerous downstream AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。