通过显式建模说话人一致性,提升语音分离效果
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
- 引入基于中心点的说话人一致性损失,增强语音一致性
- 在多个数据集上实现显著性能提升,尤其在低信噪比下表现更优
- 适合需要精准提取特定说话人语音的应用场景
目标说话人提取(TSE)利用参考声纹从混合语音中分离出目标说话人。现有基于音频线索的TSE系统依赖于登记语音提取的说话人嵌入,但该嵌入易受说话人身份混淆影响。不同于以往聚焦于优化嵌入提取的方法,本文从说话人一致性角度出发,提出一种考虑说话人一致性的目标说话人提取方法,引入基于中心点的说话人一致性损失,以确保登记语音与提取语音间的说话人一致性。此外,训练过程中还整合了条件性损失抑制机制。实验结果验证了该方法在提升TSE性能方面的有效性。语音演示可在线查看:https://sc-tse.netlify.app/
原文摘要 · Abstract (English)
Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may suffer from speaker identity confusion. Unlike previous studies that focus on improving speaker embedding extraction, we improve TSE performance from the perspective of speaker consistency. In this paper, we propose a speaker consistency-aware target speaker extraction method that incorporates a centroid-based speaker consistency loss. This approach enhances TSE performance by ensuring speaker consistency between the enrolled and extracted speech. In addition, we integrate conditional loss suppression into the training process. The experimental results validate the effectiveness of our proposed methods in advancing the TSE performance. A speech demo is available online:https://sc-tse.netlify.app/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。