用3D骨骼与手型信息提升手语边界检测精度,助力连续手语识别。
MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition
- 融合3D骨骼动态与87类手型特征,定位手语起止帧
- 在ASLLRP数据集上边界检测错误率降低12.6%,识别准确率提升5.3%
- 适合手语识别、无障碍交互与多模态动作分析研究者
本文提出一种多模态手语边界检测方法,用于连续美国手语(ASL)识别。首先利用3D骨骼特征捕捉手语动作的动态特性,这些特征在手语边界处趋于聚集。其次,通过预训练的手型分类器识别87种语言学定义的典型手型,以辅助判断手语的起止位置。采用多模态融合模块整合视频分割框架与手型分类模型。最终,使用包含孤立手势和人工标注连续手语片段的大规模数据库训练识别模型。在ASLLRP数据集上的实验表明,本方法显著优于先前工作,边界检测误差降低12.6%,手语识别准确率提升5.3%。
原文摘要 · Abstract (English)
This paper employs a multimodal approach for continuous sign recognition by first using ML for detecting the start and end frames of signs in videos of American Sign Language (ASL) sentences, and then by recognizing the segmented signs. For improved robustness we use 3D skeletal features extracted from sign language videos to take into account the convergence of sign properties and their dynamics that tend to cluster at sign boundaries. Another focus of this paper is the incorporation of information from 3D handshape for boundary detection. To detect handshapes normally expected at the beginning and end of signs, we pretrain a handshape classifier for detection of 87 linguistically defined canonical handshape categories using a dataset that we created by integrating and normalizing several existing datasets. A multimodal fusion module is then used to unify the pretrained sign video segmentation framework and handshape classification models. Finally, the estimated boundaries are used for sign recognition, where the recognition model is trained on a large database containing both citation-form isolated signs and signs pre-segmented (based on manual annotations) from continuous signing-as such signs often differ a bit in certain respects. We evaluate our method on the ASLLRP corpus and demonstrate significant improvements over previous work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。