arXiv:2507.10755cs.CVcs.AI2025-07

审计两大面部表情数据集,发现多为摆拍且存在种族偏见。

Auditing Facial Emotion Recognition Datasets for Posed Expressions and Racial Bias

  • 通过新方法识别图像是否为摆拍,发现自称‘自然场景’的数据集实则多为摆拍。
  • 模型对非白人及深肤色者易误判为愤怒或悲伤,即便其在微笑。
  • 研究揭示数据集缺陷,警示开发者关注真实场景与公平性。

面部表情识别(FER)算法将人脸表情分类为快乐、悲伤或愤怒等情绪。当前面临的评估挑战是:模型在检测自发表情时性能明显低于摆拍表情。伦理与评估挑战在于,这些模型对某些种族和肤色人群表现较差。这与数据集构建过程中的采集方式密切相关。本研究审计了两个前沿的FER数据集,随机抽样并分析图像是否为自发或摆拍,提出一种识别方法。结果显示,大量图像实为摆拍,尽管数据集宣称包含自然场景图像。由于模型性能在自发与摆拍图像间差异显著,基于此类数据训练的模型在真实部署中表现不可靠。此外,我们观察样本中个体的肤色,并测试三个在各数据集上训练的模型对不同种族和肤色人群的表情预测能力。结果发现,被标记为非白人或深肤色者即使在微笑,模型更倾向于将其判断为愤怒或悲伤。这种偏见可能导致模型在实际应用中加剧社会伤害。

原文摘要 · Abstract (English)

Facial expression recognition (FER) algorithms classify facial expressions into emotions such as happy, sad, or angry. An evaluative challenge facing FER algorithms is the fall in performance when detecting spontaneous expressions compared to posed expressions. An ethical (and evaluative) challenge facing FER algorithms is that they tend to perform poorly for people of some races and skin colors. These challenges are linked to the data collection practices employed in the creation of FER datasets. In this study, we audit two state-of-the-art FER datasets. We take random samples from each dataset and examine whether images are spontaneous or posed. In doing so, we propose a methodology for identifying spontaneous or posed images. We discover a significant number of images that were posed in the datasets purporting to consist of in-the-wild images. Since performance of FER models vary between spontaneous and posed images, the performance of models trained on these datasets will not represent the true performance if such models were to be deployed in in-the-wild applications. We also observe the skin color of individuals in the samples, and test three models trained on each of the datasets to predict facial expressions of people from various races and skin tones. We find that the FER models audited were more likely to predict people labeled as not white or determined to have dark skin as showing a negative emotion such as anger or sadness even when they were smiling. This bias makes such models prone to perpetuate harm in real life applications.

面部识别数据偏见伦理审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。