用视觉语音识别知识提升德语手语口动识别准确率
Transfer Learning from Visual Speech Recognition to Mouthing Recognition in German Sign Language
- 从视觉语音识别迁移知识到手语口动识别
- 多任务学习使口动识别与语音识别精度均提升
- 适合手语识别数据少时的跨任务知识迁移
手语识别系统多关注手势,但非手势特征如口动仍具重要语言信息。本文直接将口动实例分类为对应口语词汇,并探索从视觉语音识别(VSR)向德语手语口动识别迁移知识的可行性。利用三个VSR数据集:一个英文、一个德语但词汇无关、一个德语且目标词与口动数据集一致,研究任务相似性的影响。结果表明,多任务学习同时提升了口动识别与VSR的准确率及模型鲁棒性,说明口动识别应视为与VSR相关但独立的任务。本研究推动了在口动标注有限情况下,从VSR向SLR的知识迁移。
原文摘要 · Abstract (English)
Sign Language Recognition (SLR) systems primarily focus on manual gestures, but non-manual features such as mouth movements, specifically mouthing, provide valuable linguistic information. This work directly classifies mouthing instances to their corresponding words in the spoken language while exploring the potential of transfer learning from Visual Speech Recognition (VSR) to mouthing recognition in German Sign Language. We leverage three VSR datasets: one in English, one in German with unrelated words and one in German containing the same target words as the mouthing dataset, to investigate the impact of task similarity in this setting. Our results demonstrate that multi-task learning improves both mouthing recognition and VSR accuracy as well as model robustness, suggesting that mouthing recognition should be treated as a distinct but related task to VSR. This research contributes to the field of SLR by proposing knowledge transfer from VSR to SLR datasets with limited mouthing annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。