解决人脸识别中老年人表情识别偏差问题
Bridging the gap in FER: addressing age bias in deep learning
- 通过可解释AI分析模型对不同年龄表情的关注差异
- 三种改进策略使老年人表情识别准确率显著提升
- 仅用粗略年龄标签即可有效改善公平性
基于深度学习的面部表情识别(FER)系统近年来表现优异,但常存在年龄相关的群体偏差,尤其影响老年人群的识别可靠性。本文针对老年群体开展系统研究,分析不同年龄组的表情识别性能差异、受偏见影响最严重的表情类型,以及模型注意力模式的变化。结合可解释AI技术发现,'中性'、'悲伤'和'愤怒'在老年人中的识别误差较大,且模型关注区域存在系统性偏差。基于此,提出多任务学习、多模态输入和年龄加权损失三种偏差缓解策略。模型在大规模数据集AffectNet上训练,使用自动估算的年龄标签,并在包含少数群体的平衡基准数据集上验证。结果表明,老年人群的表情识别准确率普遍提高,尤其是高错误率表情。显著性热图分析显示,采用年龄感知策略的模型能更合理地关注各年龄段的关键面部区域,解释了性能提升原因。研究证明,仅通过简单的训练优化即可有效缓解年龄偏差,且粗略年龄标签在大规模情感计算系统中具有重要价值。
原文摘要 · Abstract (English)
Facial Expression Recognition (FER) systems based on deep learning have achieved impressive performance in recent years. However, these models often exhibit demographic biases, particularly with respect to age, which can compromise their fairness and reliability. In this work, we present a comprehensive study of age-related bias in deep FER models, with a particular focus on the elderly population. We first investigate whether recognition performance varies across age groups, which expressions are most affected, and whether model attention differs depending on age. Using Explainable AI (XAI) techniques, we identify systematic disparities in expression recognition and attention patterns, especially for "neutral", "sadness", and "anger" in elderly individuals. Based on these findings, we propose and evaluate three bias mitigation strategies: Multi-task Learning, Multi-modal Input, and Age-weighted Loss. Our models are trained on a large-scale dataset, AffectNet, with automatically estimated age labels and validated on balanced benchmark datasets that include underrepresented age groups. Results show consistent improvements in recognition accuracy for elderly individuals, particularly for the most error-prone expressions. Saliency heatmap analysis reveals that models trained with age-aware strategies attend to more relevant facial regions for each age group, helping to explain the observed improvements. These findings suggest that age-related bias in FER can be effectively mitigated using simple training modifications, and that even approximate demographic labels can be valuable for promoting fairness in large-scale affective computing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。