arXiv:2502.11049cs.CV2025-02被引 16

系统评估面部表情识别数据与模型的偏见,发现高准确率不等于公平。

Faces of Fairness: Examining Bias in Facial Expression Recognition Datasets and Models

  • 构建统一框架,整合七项指标量化数据与模型偏见。
  • 四大数据集均存显著种族偏见,AffectNet最严重,Fer2013最轻。
  • 基于Transformer的模型虽准确但偏见最高,需联合优化数据与模型。

自动面部表情识别(FER)涉及数据与模型设计两大关键环节,二者均显著影响任务中的偏见与公平性。然而,现有研究对FER数据集与模型的偏见问题仍关注不足。本研究系统分析了四种主流在野数据集(AffectNet、ExpW、Fer2013、RAF-DB)的偏见,并评估了七种深度模型(MobileNet、ResNet、XceptionNet、ViT、CLIP、POSTER、CEPrompt)的公平性。不同于以往仅考察单一偏见维度的研究,本文提出一个统一评估框架,整合五项已有及两项新提出的数据集度量指标,并结合四项模型公平性准则。此外,引入两个新指标——条件熵偏见指数与集中度指数,用于捕捉现有方法未能涵盖的条件依赖关系与组内数据不平衡。结果表明,所有数据集均存在显著的人口统计学偏见,尤以种族为甚;其中AffectNet整体偏见最高,Fer2013最低。模型层面,残差结构的CNN(ResNet、XceptionNet)偏见最低,而基于Transformer的模型(ViT、CLIP)偏见最高,尽管其准确率常达或超过其他模型。这说明高准确率不等同于公平性,必须协同优化数据与模型以实现真正公平的表达识别。

原文摘要 · Abstract (English)

Automated Facial Expression Recognition (FER), involves two critical aspects: data and model design. Both significantly influence bias and fairness in FER tasks. However, issues related to bias and fairness in FER datasets and models remain underexplored. This study investigates bias and fairness in FER datasets and models. The bias of four common in-the-wild FER datasets, including AffectNet, ExpW, Fer2013, and RAF-DB, is studied. Additionally, this research evaluates the bias and fairness of seven deep models, including three generic CNN models: MobileNet, ResNet, XceptionNet, as well as two popular Transformer-based models: ViT and CLIP, plus two FER-specific state-of-the-art models: POSTER and CEPrompt. Unlike prior studies that examine only limited aspects of bias, our work introduces a unified evaluation framework for FER that integrates five existing and two newly proposed dataset metrics with four fairness criteria for model analysis. We further introduce two new metrics, Conditional-Entropy Bias Index and Concentration Index, designed to quantify conditional dependencies and intra-group data imbalance that existing measures fail to capture. Our results show that all four datasets carry significant demographic bias, most notably in race, with AffectNet exhibiting the highest overall bias and Fer2013 the lowest. At the model level, we find that residual-based CNN architectures (ResNet and XceptionNet) exhibit the lowest overall bias, whereas Transformer-based models (ViT and CLIP) exhibit the highest, despite often achieving comparable or superior accuracy. These findings demonstrate that high predictive accuracy does not guarantee fairness, and that dataset-level and model-level bias must be addressed jointly rather than in isolation.

面部识别公平性偏见分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。