提出全局-局部融合的专家路由机制,提升重叠语音识别准确率。
GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
- 动态融合说话人全局上下文与局部声学特征,指导专家选择。
- 在LibriSpeechMix和CH109数据集上显著优于现有SOT方法。
- 特别适合高重叠复杂场景,首次将全局-局部MoE用于多说话人识别。
端到端多说话人自动语音识别(MTASR)在处理重叠语音时面临巨大挑战,关键瓶颈在于说话人特有的声学特征在深层网络中常被稀释。为此,本文提出全局-局部感知动态专家混合(GLAD)架构。GLAD引入新型路由机制,动态融合说话人感知的全局上下文与精细局部声学细节,自适应地引导专家选择。在LibriSpeechMix和CH109数据集上的实验表明,GLAD显著优于现有的基于序列输出训练(SOT)的MTASR方法,在高重叠复杂场景下表现出卓越鲁棒性。据我们所知,这是首个将全局-局部融合的MoE策略应用于MTASR的工作。
原文摘要 · Abstract (English)
End-to-end multi-talker automatic speech recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech. A critical bottleneck is that speaker-specific acoustic characteristics, which are essential for distinguishing overlapping speech, are often diluted in deep network layers. To address this, we propose the Global-Local Aware Dynamic Mixture-of-Experts (GLAD) architecture. GLAD introduces a novel routing mechanism that dynamically fuses speaker-aware global context with fine-grained local acoustic details to adaptively guide expert selection. Experiments on the LibriSpeechMix and CH109 datasets demonstrate that GLAD significantly outperforms existing Serialized Output Training (SOT)-based MTASR approaches, exhibiting exceptional robustness in challenging, high-overlap scenarios. To the best of our knowledge, this is the first work to apply a global-local fusion MoE strategy to MTASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。