用真实录音生成伪标签,让语音增强模型更适应远场环境。
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
- 通过直接声估计获取真实录音的伪标签,实现无需真实标注的训练。
- 在MISP2023数据集上显著超越现有最优方法,提升语音识别准确率。
- 适合需要部署于真实远场场景的语音增强系统研发者使用。
由于真实远场对话数据集缺乏目标语音标注,语音增强(SE)模型通常在模拟数据上训练,但在实际环境中表现不佳。为此,本文提出直接声估计(DSE),用于估计真实录音中的理想直达声,并设计一种新型伪监督学习方法SuPseudo,利用DSE生成的伪标签,使SE模型能直接从真实数据中学习并适应,从而提升泛化能力。同时,构建了专用的FARNET模型以充分融合SuPseudo。在MISP2023语料库上的实验表明,该方法有效,系统性能显著优于当前最先进水平。演示地址:https://EeLLJ.github.io/SuPseudo/
原文摘要 · Abstract (English)
Due to the lack of target speech annotations in real-recorded far-field conversational datasets, speech enhancement (SE) models are typically trained on simulated data. However, the trained models often perform poorly in real-world conditions, hindering their application in far-field speech recognition. To address the issue, we (a) propose direct sound estimation (DSE) to estimate the oracle direct sound of real-recorded data for SE; and (b) present a novel pseudo-supervised learning method, SuPseudo, which leverages DSE-estimates as pseudo-labels and enables SE models to directly learn from and adapt to real-recorded data, thereby improving their generalization capability. Furthermore, an SE model called FARNET is designed to fully utilize SuPseudo. Experiments on the MISP2023 corpus demonstrate the effectiveness of SuPseudo, and our system significantly outperforms the previous state-of-the-art. A demo of our method can be found at https://EeLLJ.github.io/SuPseudo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。