arXiv:2502.10447eess.AScs.CL2025-02ICML被引 6

用分层专家机制提升音视频语音识别的鲁棒性与扩展性

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

  • 采用稀疏专家混合架构,按输入动态激活音视频专家组
  • 在LRS3和MuAViC上实现新纪录性能,计算开销低
  • 适合需要高鲁棒性且资源受限的实时语音识别场景

音视频语音识别(AVSR)通过融合听觉与视觉信息,在嘈杂环境中显著提升识别效果。然而现有系统难以在保持计算效率的同时扩展模型规模。本文提出MoHAVE(分层音视频专家混合模型),一种新型鲁棒AVSR框架。其基于稀疏专家混合(MoE)架构,动态激活特定模态的专家组,实现对不同音视频输入的自适应处理,同时维持低计算开销。关键贡献包括:(1)可高效扩展模型容量的稀疏MoE框架;(2)基于输入上下文动态调用专家组的分层门控机制,增强适应性与鲁棒性;(3)在鲁棒AVSR基准数据集LRS3和MuAViC的转录与翻译任务中表现卓越,树立了可扩展语音识别的新标准。

原文摘要 · Abstract (English)

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up without compromising computational efficiency. In this study, we introduce MoHAVE (Mixture of Hierarchical Audio-Visual Experts), a novel robust AVSR framework designed to address these scalability constraints. By leveraging a Mixture-of-Experts (MoE) architecture, MoHAVE activates modality-specific expert groups, ensuring dynamic adaptation to various audio-visual inputs with minimal computational overhead. Key contributions of MoHAVE include: (1) a sparse MoE framework that efficiently scales AVSR model capacity, (2) a hierarchical gating mechanism that dynamically utilizes the expert groups based on input context, enhancing adaptability and robustness, and (3) remarkable performance across robust AVSR benchmarks, including LRS3 and MuAViC transcription and translation tasks, setting a new standard for scalable speech recognition systems.

音视频识别专家混合鲁棒识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。