arXiv:2607.01566cs.SD2026-07

通过显式声学引导提升多说话人语音识别的专家协作能力

H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR

论文配图:H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR
图 1 · 摘自论文原文
  • 引入说话人感知全局编码器与重叠感知损失,捕捉长程依赖
  • 在LibriSpeechMix上比基线提升显著,复杂重叠场景效果更优
  • 适合研究多说话人语音识别与MoE模型融合的学者参考

多说话人自动语音识别(MTASR)在复杂重叠语音场景下面临巨大挑战。现有混合专家(MoE)方法通常依赖帧无关路由,导致时间感知局限,并仅通过下游ASR目标进行隐式表征学习,缺乏明确指导。为此,本文提出面向多说话人语音识别的全貌说话人感知引导专家(H-SAGE)。具体地,设计说话人感知全局编码器,结合辅助的重叠感知损失,显式引导模型区分不同声学状态;同时提出全貌门控机制,综合评估全局上下文与局部细节以决策专家选择。在LibriSpeechMix数据集上的实验表明,H-SAGE在复杂重叠场景下持续优于强基线,验证了显式声学引导能有效增强专家协作。代码已公开于https://github.com/NKU-HLT/H-SAGE。

原文摘要 · Abstract (English)

Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. While recent Mixture-of-Experts (MoE) approaches have shown promise, they typically rely on frame-independent routing that leads to temporal myopia, and depend solely on the downstream ASR objective, which results in implicit and ungrounded representation learning. To address these limitations, we propose Holistic Speaker-Aware Guided Experts (H-SAGE) for MoE-based MTASR. Specifically, we introduce a Speaker-Aware Global Encoder to capture long-term dependencies, supervised by an auxiliary Overlap-Aware Loss that explicitly guides the model to discern acoustic states. Furthermore, we design a Holistic Gating Mechanism to arbitrate expert selection by jointly evaluating global context and local details. Experiments on LibriSpeechMix demonstrate that H-SAGE achieves consistent improvements over strong baselines, particularly in complex scenarios, validating that explicit acoustic guidance effectively enhances expert collaboration. Our code can be found at https://github.com/NKU-HLT/H-SAGE.

多说话人识别MoE模型语音分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。