用专家混合模型提升视频中语音识别的鲁棒性
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
- 采用视觉专家混合模块融合多源视觉信息
- 在三个基准上达到当前最优性能
- 适合处理真实场景下复杂视频的语音识别
视觉信号可通过提供额外上下文信息提升音视频语音识别准确率。鉴于视觉信号的复杂性,音视频语音识别模型需具备在多样化视频场景下的强泛化能力,面临重大挑战。本文提出EVA,利用专家混合(Mixture-of-Experts)机制构建音频视觉语音识别模型,实现对真实环境中视频的鲁棒语音识别。具体地,首先将视觉信息编码为视觉标记序列,并通过轻量级投影映射至语音空间;随后基于一个鲁棒预训练语音识别模型构建EVA,确保其泛化能力。此外,为有效融合视觉信息,通过专家混合模块将视觉信号注入语音识别模型。实验表明,该模型在三个基准测试上均取得领先结果,验证了EVA在多种视频领域中的优异泛化能力。
原文摘要 · Abstract (English)
Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities across diverse video scenarios, presenting a significant challenge. In this paper, we introduce EVA, leveraging the mixture-of-Experts for audioVisual ASR to perform robust speech recognition for ``in-the-wild'' videos. Specifically, we first encode visual information into visual tokens sequence and map them into speech space by a lightweight projection. Then, we build EVA upon a robust pretrained speech recognition model, ensuring its generalization ability. Moreover, to incorporate visual information effectively, we inject visual information into the ASR model through a mixture-of-experts module. Experiments show our model achieves state-of-the-art results on three benchmarks, which demonstrates the generalization ability of EVA across diverse video domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。