arXiv:2411.13811cs.SDcs.MM2024-11被引 1

X-CrossNet通过交叉注意力融合目标说话人特征,提升嘈杂环境下的语音分离效果。

X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion

  • 采用CrossNet骨干网络,专为噪声和混响环境优化。
  • 在WSJ0-2mix和WHAMR!数据集上性能超越现有方法。
  • 适合需要高鲁棒性语音提取的实际场景应用。

目标说话人提取(TSE)是一种利用目标说话人的辅助特征从混合语音中分离出其声音的技术,是解决鸡尾酒会问题的新型尝试,具有比传统语音分离更广阔的应用前景。尽管学术界在此领域已取得优异的公开数据集表现,但多数模型在真实噪声或混响环境下性能显著下降。为此,本文提出新型TSE模型X-CrossNet,以专门针对复杂声学环境优化的CrossNet为骨干网络。为进一步增强模型对目标说话人辅助特征的捕捉与利用能力,我们在每个CrossNet模块的全局多头自注意力(GMHSA)中引入交叉注意力机制,实现目标说话人特征与混合语音特征的更高效融合。实验结果表明,该方法在WSJ0-2mix和WHAMR!数据集上均取得更优分离效果,展现出强鲁棒性和稳定性。

原文摘要 · Abstract (English)

Target speaker extraction (TSE) is a technique for isolating a target speaker's voice from mixed speech using auxiliary features associated with the target speaker. It is another attempt at addressing the cocktail party problem and is generally considered to have more practical application prospects than traditional speech separation methods. Although academic research in this area has achieved high performance and evaluation scores on public datasets, most models exhibit significantly reduced performance in real-world noisy or reverberant conditions. To address this limitation, we propose a novel TSE model, X-CrossNet, which leverages CrossNet as its backbone. CrossNet is a speech separation network specifically optimized for challenging noisy and reverberant environments, achieving state-of-the-art performance in tasks such as speaker separation under these conditions. Additionally, to enhance the network's ability to capture and utilize auxiliary features of the target speaker, we integrate a Cross-Attention mechanism into the global multi-head self-attention (GMHSA) module within each CrossNet block. This facilitates more effective integration of target speaker features with mixed speech features. Experimental results show that our method performs superior separation on the WSJ0-2mix and WHAMR! datasets, demonstrating strong robustness and stability.

语音分离交叉注意力鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。