arXiv:2411.03085cs.SDcs.LG2024-11被引 15

用预训练前端缩小真实与合成语音分离的差距。

Speech Separation with Pretrained Frontend to Minimize Domain Mismatch

  • 设计自监督前端,无参考语音也能学习混合信号共性特征。
  • 在真实数据上分离效果提升,相较基线显著改善信噪比。
  • 适合需在复杂现实场景中部署语音分离系统的开发者。

语音分离旨在从混合语音中分离出个体说话人信号。由于真实场景下缺乏目标参考语音,多数模型在合成数据上训练,导致真实与合成数据间存在领域差异。本文提出一种自监督域不变预训练(DIP)前端,无需目标语音即可处理混合数据。该前端采用孪生网络,结合混合预测编码(MPC)和混合不变编码(MIC)两项创新预训练任务,捕捉真实与合成未标注混合信号间的共享上下文线索。随后将DIP前端冻结为特征提取器,在合成数据上训练下游分离模型。通过预训练获取上下文信息,使合成数据上学到的分离能力能有效迁移到真实数据。为更好利用该前端,我们设计新分离流水线以对齐特征分辨率。在标准基准与真实数据集上的评估表明,DIP前端优于现有模型,验证了大规模预训练在提升真实场景语音分离质量与可懂度方面的潜力。

原文摘要 · Abstract (English)

Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail party scenarios. As a result, there exists a domain gap between real and synthetic data when deploying speech separation models in real-world applications. In this paper, we propose a self-supervised domain-invariant pretrained (DIP) frontend that is exposed to mixture data without the need for target reference speech. The DIP frontend utilizes a Siamese network with two innovative pretext tasks, mixture predictive coding (MPC) and mixture invariant coding (MIC), to capture shared contextual cues between real and synthetic unlabeled mixtures. Subsequently, we freeze the DIP frontend as a feature extractor when training the downstream speech separation models on synthetic data. By pretraining the DIP frontend with the contextual cues, we expect that the speech separation skills learned from synthetic data can be effectively transferred to real data. To benefit from the DIP frontend, we introduce a novel separation pipeline to align the feature resolution of the separation models. We evaluate the speech separation quality on standard benchmarks and real-world datasets. The results confirm the superiority of our DIP frontend over existing speech separation models. This study underscores the potential of large-scale pretraining to enhance the quality and intelligibility of speech separation in real-world applications.

语音分离预训练域适应自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。