arXiv:2506.14750cs.SDcs.AI2025-06被引 1

用专家混合机制提升语音分说话人系统在复杂环境下的表现

Exploring Speaker Diarization with Mixture of Experts

  • 引入记忆感知的多说话人嵌入与序列到序列框架
  • 在多个真实场景数据集上实现当前最优性能
  • 适合需要高鲁棒性语音分析的应用开发者

本文提出一种基于记忆感知多说话人嵌入与序列到序列架构的新型神经说话人分组系统(NSD-MS2S),通过记忆模块增强说话人嵌入,并利用Seq2Seq框架高效映射声学特征至说话人标签。进一步探索专家混合在说话人分组中的应用,提出共享软专家混合(SS-MoE)模块以缓解模型偏差并提升性能,形成扩展模型NSD-MS2S-SSMoE。在CHiME-6、DiPCo、Mixer 6和DIHARD-III等多个复杂声学数据集上的实验表明,该方法显著提升了系统的鲁棒性与泛化能力,达到当前最佳性能,验证了其在真实复杂场景中的有效性。

原文摘要 · Abstract (English)

In this paper, we propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates a memory-aware multi-speaker embedding module with a sequence-to-sequence architecture. The system leverages a memory module to enhance speaker embeddings and employs a Seq2Seq framework to efficiently map acoustic features to speaker labels. Additionally, we explore the application of mixture of experts in speaker diarization, and introduce a Shared and Soft Mixture of Experts (SS-MoE) module to further mitigate model bias and enhance performance. Incorporating SS-MoE leads to the extended model NSD-MS2S-SSMoE. Experiments on multiple complex acoustic datasets, including CHiME-6, DiPCo, Mixer 6 and DIHARD-III evaluation sets, demonstrate meaningful improvements in robustness and generalization. The proposed methods achieve state-of-the-art results, showcasing their effectiveness in challenging real-world scenarios.

说话人分组专家混合语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。