arXiv:2410.16059eess.AScs.SD2024-10被引 24

多层级语音表征提升目标说话人分离效果

Multi-Level Speaker Representation for Target Speaker Extraction

  • 从原始频谱到神经嵌入,融合多层级说话人特征
  • 在Libri2mix上提升2.74 dB信噪比和4.94%准确率
  • 适合需要高精度说话人分离的语音增强任务

目标说话人提取(TSE)依赖目标说话人的参考线索从语音混合中分离出目标语音。虽然常用说话人嵌入作为参考线索,但基于大量说话人预训练的嵌入可能产生身份混淆。本文提出一种多层级说话人表征方法,涵盖原始特征到神经嵌入,作为说话人参考线索。我们从登记语音的幅度谱图生成谱级表征,作为低层原始特征,显著提升模型泛化能力。此外,提出基于交叉注意力机制的上下文嵌入特征,融合预训练说话人编码器的帧级嵌入。通过整合多层级说话人特征,显著提升TSE模型性能。在Libri2mix测试集上,相比基线模型,信噪比提升2.74 dB,提取准确率提高4.94%。

原文摘要 · Abstract (English)

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In this work, we propose a multi-level speaker representation approach, from raw features to neural embeddings, to serve as the speaker reference cue. We generate a spectral-level representation from the enrollment magnitude spectrogram as a raw, low-level feature, which significantly improves the model's generalization capability. Additionally, we propose a contextual embedding feature based on cross-attention mechanisms that integrate frame-level embeddings from a pre-trained speaker encoder. By incorporating speaker features across multiple levels, we significantly enhance the performance of the TSE model. Our approach achieves a 2.74 dB improvement and a 4.94% increase in extraction accuracy on Libri2mix test set over the baseline.

说话人分离多层级表征语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。