通过引导视觉发音特征,提升语音识别在噪声下的鲁棒性。
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
- 用附加的发音部位分类任务,让模型更关注视觉发音信息。
- 在-10dB信噪比下字错误率从13.59%降至6.60%,相对提升51.4%。
- 适合研究音视频语音识别、噪声环境优化的开发者参考。
当前音视频语音识别(AVSR)系统多采用大语言模型解码器与基于Transformer的编码器结合,达到顶尖性能。然而,性能提升是来自语言建模改进,还是音视频编码增强尚不明确。本文提出Viseme-Guided AV-HuBERT(VisG AV-HuBERT),一种多任务微调框架,通过引入辅助发音部位(viseme)分类任务,强化模型对视觉发音特征的依赖。在原AV-HuBERT基础上增加轻量级发音部位预测子网络,显式引导编码器保留视觉语音信息。在LRS3数据集上评估,该方法表现与基线相当或更优,尤其在高噪声条件下效果显著:在语音噪声环境下,字错误率(WER)从13.59%降至6.60%(相对降低51.4%)。深入分析显示,各类噪声下替换错误大幅减少,表明语音单元区分能力增强。在LRS2上的测试验证了其泛化能力。结果表明,显式建模发音部位可改善编码器表征,为通过编码器层面优化实现噪声鲁棒的音视频语音识别提供基础。
原文摘要 · Abstract (English)
Audio-Visual Speech Recognition (AVSR) systems nowadays integrate Large Language Model (LLM) decoders with transformer-based encoders, achieving state-of-the-art results. However, the relative contributions of improved language modelling versus enhanced audiovisual encoding remain unclear. We propose Viseme-Guided AV-HuBERT (VisG AV-HuBERT), a multi-task fine-tuning framework that incorporates auxiliary viseme classification to strengthen the model's reliance on visual articulatory features. By extending AV-HuBERT with a lightweight viseme prediction sub-network, this method explicitly guides the encoder to preserve visual speech information. Evaluated on LRS3, VisG AV-HuBERT achieves comparable or improved performance over the baseline AV-HuBERT, with notable gains under heavy noise conditions. WER reduces from 13.59% to 6.60% (51.4% relative improvement) at -10 dB Signal-to-Noise Ratio (SNR) for Speech noise. Deeper analysis reveals substantial reductions in substitution errors across noise types, demonstrating improved speech unit discrimination. Evaluation on LRS2 confirms generalization capability. Our results demonstrate that explicit viseme modelling enhances encoder representations, and provides a foundation for enhancing noise-robust AVSR through encoder-level improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。