用面部关键点指导特征提取,提升少样本下语音识别准确率
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
- 以面部关键点为辅助信息,引导视觉特征提取
- 在有限数据下实现更高识别准确率,对未见说话人也有效
- 适合资源受限场景下的视觉语音识别任务
视觉语音识别旨在从无声视频中识别口语内容,近年来受到广泛关注。深度学习方法虽显著提升了识别速度与精度,但仍易受光照、皮肤纹理等用户特异性因素干扰。现有数据驱动方法依赖大规模预训练模型缓解此类问题,但需大量数据与计算资源。为此,本文提出一种关键点引导的视觉特征提取器,利用面部关键点作为辅助信息,设计时空多图卷积网络,充分挖掘关键点的空间位置与时序动态特征,并引入多层级唇部动态融合框架,将关键点特征与原始视频帧特征结合。实验表明,该方法在数据有限条件下表现优异,且能提升对未见说话人的识别准确率。
原文摘要 · Abstract (English)
Visual speech recognition is a technique to identify spoken content in silent speech videos, which has raised significant attention in recent years. Advancements in data-driven deep learning methods have significantly improved both the speed and accuracy of recognition. However, these deep learning methods can be effected by visual disturbances, such as lightning conditions, skin texture and other user-specific features. Data-driven approaches could reduce the performance degradation caused by these visual disturbances using models pretrained on large-scale datasets. But these methods often require large amounts of training data and computational resources, making them costly. To reduce the influence of user-specific features and enhance performance with limited data, this paper proposed a landmark guided visual feature extractor. Facial landmarks are used as auxiliary information to aid in training the visual feature extractor. A spatio-temporal multi-graph convolutional network is designed to fully exploit the spatial locations and spatio-temporal features of facial landmarks. Additionally, a multi-level lip dynamic fusion framework is introduced to combine the spatio-temporal features of the landmarks with the visual features extracted from the raw video frames. Experimental results show that this approach performs well with limited data and also improves the model's accuracy on unseen speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。