arXiv:2409.02615eess.AScs.SD2024-09中稿 · IEEE Transactions …被引 27

无需说话人嵌入,用注意力机制实现更精准的语音分离

USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

  • 用多头交叉注意力提取目标说话人特征,不依赖说话人识别模型
  • 在多个数据集上达到最新最佳性能,尤其在噪声和混响环境下表现优异
  • 可无缝集成到各类语音分离模型中,适合实际部署场景

目标说话人提取旨在从混合语音中分离出特定说话人的声音。传统方法依赖从参考语音中提取说话人嵌入,需使用说话人识别模型,但模型选择困难且嵌入信息未必最优。本文提出通用无说话人嵌入的目标说话人提取框架(USEF-TSE),无需依赖说话人嵌入。该方法采用多头交叉注意力机制作为帧级目标说话人特征提取器,使主流说话人提取方案摆脱对说话人识别模型的依赖,并更好利用录音语音中的说话人特征与上下文信息。此外,USEF-TSE可无缝集成至时域或时频域语音分离模型,实现高效说话人提取。实验表明,该方法在WSJ0-2mix、WHAM!、WHAMR!等标准基准上,于单声道无回声、噪声及噪声混响双说话人分离与提取任务中,以尺度无关信噪比(SI-SDR)指标达到当前最佳性能。在LibriMix及ICASSP 2023 DNS挑战赛盲测集上的结果也显示,模型在更多样和域外数据上表现稳健。源代码请访问:https://github.com/ZBang/USEF-TSE。

原文摘要 · Abstract (English)

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recognition model is required. However, identifying an appropriate speaker recognition model can be challenging, and using the target speaker embedding as reference information may not be optimal for target speaker extraction tasks. This paper introduces a Universal Speaker Embedding-Free Target Speaker Extraction (USEF-TSE) framework that operates without relying on speaker embeddings. USEF-TSE utilizes a multi-head cross-attention mechanism as a frame-level target speaker feature extractor. This innovative approach allows mainstream speaker extraction solutions to bypass the dependency on speaker recognition models and better leverage the information available in the enrollment speech, including speaker characteristics and contextual details. Additionally, USEF-TSE can seamlessly integrate with other time-domain or time-frequency domain speech separation models to achieve effective speaker extraction. Experimental results show that our proposed method achieves state-of-the-art (SOTA) performance in terms of Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) on the WSJ0-2mix, WHAM!, and WHAMR! datasets, which are standard benchmarks for monaural anechoic, noisy and noisy-reverberant two-speaker speech separation and speaker extraction. The results on the LibriMix and the blind test set of the ICASSP 2023 DNS Challenge demonstrate that the model performs well on more diverse and out-of-domain data. For access to the source code, please visit: https://github.com/ZBang/USEF-TSE.

说话人分离注意力机制语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。