改进ConvNeXt模型,提升面部表情识别准确率
EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
- 在ConvNeXt基础上引入空间变换网络与注意力机制
- 在FER2013数据集上达到更高分类准确率
- 适合需要高精度表情识别的AI应用开发者
面部表情在人类交流中至关重要,是表达多种情绪的重要方式。随着人工智能与计算机视觉的发展,深度神经网络已成为面部表情识别的有效工具。本文提出EmoNeXt,一种基于改进ConvNeXt架构的新型深度学习框架。通过集成空间变换网络(STN)聚焦面部特征丰富区域,并引入压缩-激励模块捕捉通道间依赖关系。此外,设计自注意力正则化项,促使模型生成更紧凑的特征向量。实验表明,该模型在FER2013数据集上的情绪分类性能优于现有先进深度学习模型。
原文摘要 · Abstract (English)
Facial expressions play a crucial role in human communication serving as a powerful and impactful means to express a wide range of emotions. With advancements in artificial intelligence and computer vision, deep neural networks have emerged as effective tools for facial emotion recognition. In this paper, we propose EmoNeXt, a novel deep learning framework for facial expression recognition based on an adapted ConvNeXt architecture network. We integrate a Spatial Transformer Network (STN) to focus on feature-rich regions of the face and Squeeze-and-Excitation blocks to capture channel-wise dependencies. Moreover, we introduce a self-attention regularization term, encouraging the model to generate compact feature vectors. We demonstrate the superiority of our model over existing state-of-the-art deep learning models on the FER2013 dataset regarding emotion classification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。