arXiv:2507.20884cs.CVcs.CL2025-07中稿 · 9th International …被引 4

研究发现嘴部特征对手语识别最关键,比眼睛或全脸更有效。

The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?

  • 用深度学习模型对比眼睛、嘴部和全脸区域的作用
  • 嘴部特征使识别准确率显著提升,优于其他部位
  • 适合做手语识别系统优化的研究者参考

非手动面部特征在手语交流中至关重要,但在自动手语识别(ASLR)中的作用仍缺乏深入探索。以往研究虽表明引入面部特征可提升识别效果,但多依赖手工特征提取,且仅比较仅手部特征与手部+面部特征的组合。本文通过两种深度学习模型(基于CNN和基于Transformer)在随机选取类别的孤立手势数据集上,系统分析眼睛、嘴部和全脸三个面部区域的贡献。通过定量性能评估与定性注意力图分析,结果表明嘴部是最重要的非手动面部特征,能显著提高识别准确率。研究强调了在ASLR中纳入面部特征的必要性。

原文摘要 · Abstract (English)

Non-manual facial features play a crucial role in sign language communication, yet their importance in automatic sign language recognition (ASLR) remains underexplored. While prior studies have shown that incorporating facial features can improve recognition, related work often relies on hand-crafted feature extraction and fails to go beyond the comparison of manual features versus the combination of manual and facial features. In this work, we systematically investigate the contribution of distinct facial regionseyes, mouth, and full faceusing two different deep learning models (a CNN-based model and a transformer-based model) trained on an SLR dataset of isolated signs with randomly selected classes. Through quantitative performance and qualitative saliency map evaluation, we reveal that the mouth is the most important non-manual facial feature, significantly improving accuracy. Our findings highlight the necessity of incorporating facial features in ASLR.

手语识别面部特征深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。