arXiv:2409.09589cs.SDeess.AS2024-09中稿 · SLT2024被引 16

通过增强注册语音提升说话人提取效果,最高增益2.5dB。

On the effectiveness of enrollment speech augmentation for Target Speaker Extraction

  • 在注册语音阶段引入数据增强,提升模型泛化能力
  • 新方法SSA使性能最高提升2.5dB(Libri2Mix测试集)
  • 适合资源有限场景下的说话人提取任务

深度学习显著提升了目标说话人提取(TSE)的性能。为在训练数据不足时增强算法的泛化性与鲁棒性,数据增强是常用手段。本文深入研究了对注册语音空间进行增强的有效性。实验表明,无论使用预训练还是联合优化的说话人编码器,直接增强注册语音均能带来一致的性能提升。除常见的噪声、混响添加外,本文提出一种名为自估计语音增强(SSA)的新方法。在Libri2Mix测试集上的实验结果表明,该方法可实现最高达2.5 dB的性能提升。

原文摘要 · Abstract (English)

Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a commonly adopted technique. Unlike typical data augmentation applied to speech mixtures, this work thoroughly investigates the effectiveness of augmenting the enrollment speech space. We found that for both pretrained and jointly optimized speaker encoders, directly augmenting the enrollment speech leads to consistent performance improvement. In addition to conventional methods such as noise and reverberation addition, we propose a novel augmentation method called self-estimated speech augmentation (SSA). Experimental results on the Libri2Mix test set show that our proposed method can achieve an improvement of up to 2.5 dB.

说话人提取数据增强语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。