arXiv:2502.01547eess.AScs.CV2025-02中稿 · Signal Processing …被引 6

多语言音视频降噪语音识别新模型,性能超越纯音频方案。

mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition

  • 融合Whisper与AV-HuBERT,引入模态丢弃增强跨模态融合。
  • 在9种语言的MuAViC数据集上达到最先进词错误率。
  • 噪声环境下所有语言均优于纯音频Whisper,适合多语种场景。

音视频语音识别(AVSR)结合唇动视频与音频,在噪声环境中能提升识别效果,但多数方法仅基于英语数据训练。其主要瓶颈是缺乏大规模多语言视频数据,难以从零训练模型。本文提出mWhisper-Flamingo,融合预训练音频模型Whisper与视频模型AV-HuBERT。为增强多模态融合并提升多语言噪声环境下的性能,引入解码器模态丢弃策略,使模型在配对音视频输入及独立音/视频输入下共同训练。该模型在包含9种语言的MuAViC数据集上实现最先进的词错误率(WER)。在各类噪声条件下,音视频版本的mWhisper-Flamingo在所有语言上均显著优于纯音频Whisper。

原文摘要 · Abstract (English)

Audio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation is the lack of large-scale multilingual video data, which makes it hard to train models from scratch. In this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT). To enable better multi-modal integration and improve the noisy multilingual performance, we introduce decoder modality dropout where the model is trained both on paired audio-visual inputs and separate audio/visual inputs. mWhisper-Flamingo achieves state-of-the-art WER on MuAViC, an AVSR dataset of 9 languages. Audio-visual mWhisper-Flamingo consistently outperforms audio-only Whisper on all languages in noisy conditions.

多语言音视频降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。