轻量级模型融合CNN与ViT,实时识别司机表情
Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression
- 双模型架构:CNN与视觉变换器并行提取特征
- 在KMU-FED和KDEF数据集上达到更高准确率
- 适合车载系统等实时场景,推理速度快
现有驾驶员面部表情识别(DFER)方法通常计算开销大,难以满足实时应用需求。本文提出一种基于迁移学习的双架构模型ShuffViT-DFER,巧妙结合卷积神经网络(CNN)与视觉变换器(ViT)的优势,实现高效与高精度的平衡。通过有效融合两类模型提取的特征,显著提升了模型对驾驶员面部表情的识别能力。在两个公开基准数据集KMU-FED和KDEF上的实验结果表明,该方法在保持高效性的同时,性能优于当前主流方法,具备良好的实时应用前景。
原文摘要 · Abstract (English)
Existing methods for driver facial expression recognition (DFER) are often computationally intensive, rendering them unsuitable for real-time applications. In this work, we introduce a novel transfer learning-based dual architecture, named ShuffViT-DFER, which elegantly combines computational efficiency and accuracy. This is achieved by harnessing the strengths of two lightweight and efficient models using convolutional neural network (CNN) and vision transformers (ViT). We efficiently fuse the extracted features to enhance the performance of the model in accurately recognizing the facial expressions of the driver. Our experimental results on two benchmarking and public datasets, KMU-FED and KDEF, highlight the validity of our proposed method for real-time application with superior performance when compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。