arXiv:2412.12512cs.SDeess.AS2024-12被引 7

构建多样真实噪声下的说话人提取数据集,提升语音分离鲁棒性。

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data

  • 用真实语音与合成语音混合构建新数据集,增强说话人多样性。
  • 在真实测试集上,模型性能提升1.39 dB SDR,合成数据有效改善表现。
  • 采用分阶段训练策略,适合追求高鲁棒性的语音处理研究者。

目标说话人提取(TSE)在复杂声学环境下的语音处理中至关重要。现有系统受限于数据多样性不足和真实场景下鲁棒性差,主要因训练数据为人工混叠、说话人变化少且噪声不真实。为此,我们提出Libri2Vox数据集,将来自LibriTTS的清晰目标语音与来自VoxCeleb2的干扰语音结合,构建大规模、多样的真实噪声场景数据。同时,利用先进语音生成模型合成说话人以进一步扩充多样性。为更有效融合合成数据,引入课程学习策略,按难度逐步训练模型。多架构实验表明,各模型均有提升,其中SpeakerBeam在Libri2Talker测试集上相比基线提升1.39 dB SDR。基于说话人相似性的课程学习进一步优化了Conformer模型,相较随机采样方法额外提升0.78 dB。结果表明,真实数据、合成数据与结构化训练策略协同提升TSE系统鲁棒性。

原文摘要 · Abstract (English)

Target speaker extraction (TSE) is essential in speech processing applications, particularly in scenarios with complex acoustic environments. Current TSE systems face challenges in limited data diversity and a lack of robustness in real-world conditions, primarily because they are trained on artificially mixed datasets with limited speaker variability and unrealistic noise profiles. To address these challenges, we propose Libri2Vox, a new dataset that combines clean target speech from the LibriTTS dataset with interference speech from the noisy VoxCeleb2 dataset, providing a large and diverse set of speakers under realistic noisy conditions. We also augment Libri2Vox with synthetic speakers generated using state-of-the-art speech generative models to enhance speaker diversity. Additionally, to further improve the effectiveness of incorporating synthetic data, curriculum learning is implemented to progressively train TSE models with increasing levels of difficulty. Extensive experiments across multiple TSE architectures reveal varying degrees of improvement, with SpeakerBeam demonstrating the most substantial gains: a 1.39 dB improvement in signal-to-distortion ratio (SDR) on the Libri2Talker test set compared to baseline training. Building upon these results, we further enhanced performance through our speaker similarity-based curriculum learning approach with the Conformer architecture, achieving an additional 0.78 dB improvement over conventional random sampling methods in which data samples are randomly selected from the entire dataset. These results demonstrate the complementary benefits of diverse real-world data, synthetic speaker augmentation, and structured training strategies in building robust TSE systems.

说话人提取语音合成课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。