解决人脸表情识别中的数据偏差与不平衡问题,提升真实场景下预测可靠性。
GReFEL: Geometry-Aware Reliable Facial Expression Learning under Bias and Imbalanced Data Distribution
- 基于视觉变换器和几何感知锚点的可靠性平衡模块
- 在多个数据集上显著优于现有最先进方法
- 适合需要高鲁棒性表情识别的应用场景
可靠的人脸表情学习(FEL)需有效捕捉独特表情特征,以实现真实场景中更准确、无偏的预测。然而,由于个体面部结构、动作、语调及人口统计差异导致的表情变异,当前系统仍面临挑战。有偏且不平衡的数据集进一步加剧问题,引发错误和偏见的标签。为此,我们提出GReFEL,结合视觉变换器与面部几何感知锚点式可靠性平衡模块,应对数据分布不平衡、偏差及不确定性。通过整合局部与全局数据,利用学习不同面部数据点与结构特征的锚点,调整因类内差异、类间相似性及尺度敏感性导致的误标情绪,实现全面、准确且可靠的表达预测。大量实验验证了该模型在多个数据集上的优越性能。
原文摘要 · Abstract (English)
Reliable facial expression learning (FEL) involves the effective learning of distinctive facial expression characteristics for more reliable, unbiased and accurate predictions in real-life settings. However, current systems struggle with FEL tasks because of the variance in people's facial expressions due to their unique facial structures, movements, tones, and demographics. Biased and imbalanced datasets compound this challenge, leading to wrong and biased prediction labels. To tackle these, we introduce GReFEL, leveraging Vision Transformers and a facial geometry-aware anchor-based reliability balancing module to combat imbalanced data distributions, bias, and uncertainty in facial expression learning. Integrating local and global data with anchors that learn different facial data points and structural features, our approach adjusts biased and mislabeled emotions caused by intra-class disparity, inter-class similarity, and scale sensitivity, resulting in comprehensive, accurate, and reliable facial expression predictions. Our model outperforms current state-of-the-art methodologies, as demonstrated by extensive experiments on various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。