用真实录音构建训练集,让语音分离模型在现实环境更准。
Developing an Effective Training Dataset to Enhance the Performance of AI-based Speaker Separation Systems
- 通过真实录音合成混合音频与对应语音源,生成更贴近现实的训练数据。
- 在真实混音场景下,语音分离性能提升1.65 dB(SI-SDR)。
- 适合需要部署于复杂声学环境的语音分离系统开发者。
本文针对语音分离技术虽取得进展但实际录音条件下性能下降的问题展开研究。现有模型多基于合成数据训练,这些数据由软件生成,缺乏真实录音中的噪声、混响等复杂因素。为解决此问题,我们提出一种新方法,构建包含真实混合信号及其对应各说话人原始信号的训练集。在深度学习模型上评估该数据集,相比合成数据,实现了1.65 dB的尺度不变信干比(SI-SDR)提升,显著改善了真实混音场景下的分离效果。结果表明,使用真实训练数据可有效提升模型在实际应用中的表现。
原文摘要 · Abstract (English)
This paper addresses the challenge of speaker separation, which remains an active research topic despite the promising results achieved in recent years. These results, however, often degrade in real recording conditions due to the presence of noise, echo, and other interferences. This is because neural models are typically trained on synthetic datasets consisting of mixed audio signals and their corresponding ground truths, which are generated using computer software and do not fully represent the complexities of real-world recording scenarios. The lack of realistic training sets for speaker separation remains a major hurdle, as obtaining individual sounds from mixed audio signals is a nontrivial task. To address this issue, we propose a novel method for constructing a realistic training set that includes mixture signals and corresponding ground truths for each speaker. We evaluate this dataset on a deep learning model and compare it to a synthetic dataset. We got a 1.65 dB improvement in Scale Invariant Signal to Distortion Ratio (SI-SDR) for speaker separation accuracy in realistic mixing. Our findings highlight the potential of realistic training sets for enhancing the performance of speaker separation models in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。