arXiv:2410.05074cs.CV2024-10中稿 · the AIEDM Workshop…被引 4

用改进的LSTM模型提升学生表情识别准确率

xLSTM-FER: Enhancing Student Expression Recognition with Extended Vision Long Short-Term Memory Network

  • 将图像分块后用xLSTM堆叠处理,捕捉表情时序变化
  • 在CK+、RAF-DF、FERplus数据集上优于现有方法
  • 适合高分辨率图像,可高效处理非序列图像

学生表情识别已成为评估学习体验与情绪状态的重要工具。本文提出xLSTM-FER,一种基于扩展长短期记忆网络(xLSTM)的新架构,通过先进的序列处理能力提升学生面部表情识别的准确性与效率。该模型将输入图像分割为一系列图像块,并利用xLSTM块堆叠进行处理,能够捕捉真实场景下学生面部表情的细微变化,通过学习序列内的时空关系提升识别性能。在CK+、RAF-DF和FERplus数据集上的实验表明,xLSTM-FER在标准数据集上表现优于当前最优方法。其线性计算与内存复杂度使其特别适用于高分辨率图像处理,且无需额外开销即可高效处理非序列输入如图像。

原文摘要 · Abstract (English)

Student expression recognition has become an essential tool for assessing learning experiences and emotional states. This paper introduces xLSTM-FER, a novel architecture derived from the Extended Long Short-Term Memory (xLSTM), designed to enhance the accuracy and efficiency of expression recognition through advanced sequence processing capabilities for student facial expression recognition. xLSTM-FER processes input images by segmenting them into a series of patches and leveraging a stack of xLSTM blocks to handle these patches. xLSTM-FER can capture subtle changes in real-world students' facial expressions and improve recognition accuracy by learning spatial-temporal relationships within the sequence. Experiments on CK+, RAF-DF, and FERplus demonstrate the potential of xLSTM-FER in expression recognition tasks, showing better performance compared to state-of-the-art methods on standard datasets. The linear computational and memory complexity of xLSTM-FER make it particularly suitable for handling high-resolution images. Moreover, the design of xLSTM-FER allows for efficient processing of non-sequential inputs such as images without additional computation.

表情识别序列建模xLSTM教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。