用合成语音提升说话人提取模型在相似声线下的表现
Improving curriculum learning for target speaker extraction with synthetic speakers
- 用k近邻语音转换生成多样干扰声源,增强课程学习数据
- 在高相似度场景下,多个TSE模型性能显著提升
- 适合研究语音分离与自适应训练的学者参考
目标说话人提取(TSE)旨在从复杂语音环境中分离出特定说话人的声音。当说话人特征相近时,TSE系统性能常受影响。近期研究引入课程学习(CL),让模型逐步在复杂度递增的语音样本上训练:先低相似度,后高相似度。本文提出使用基于k近邻的语音转换方法,生成多样化干扰说话人语音,并将其融入课程学习过程。实验表明,基于合成说话人的训练数据能有效提升模型能力,显著改善多个TSE系统的性能。
原文摘要 · Abstract (English)
Target speaker extraction (TSE) aims to isolate individual speaker voices from complex speech environments. The effectiveness of TSE systems is often compromised when the speaker characteristics are similar to each other. Recent research has introduced curriculum learning (CL), in which TSE models are trained incrementally on speech samples of increasing complexity. In CL training, the model is first trained on samples with low speaker similarity between the target and interference speakers, and then on samples with high speaker similarity. To further improve CL, this paper uses a $k$-nearest neighbor-based voice conversion method to simulate and generate speech of diverse interference speakers, and then uses the generated data as part of the CL. Experiments demonstrate that training data based on synthetic speakers can effectively enhance the model's capabilities and significantly improve the performance of multiple TSE systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。