用真实麦克风数据生成语音增强训练目标,提升远场语音识别效果。
Generating Training Targets for Real-World Speech Enhancement via Close-to-Distant Microphone Projection

- 通过近距麦克风信号投影生成远距录音的干净参考信号。
- 在CHiME6数据集上,性能优于当前最优的GSS方法。
- 适合做远场语音增强、语音识别任务的研究者参考。
在远场语音捕捉场景下训练神经网络进行语音增强(SE)需要配对的失真与干净语音信号。尽管这些数据常通过仿真生成,但仿真与真实录音之间的差异显著限制了SE性能。为此,我们提出近距到远距麦克风投影(C2D projection)方法,利用真实近距和远距麦克风录制的数据生成配对数据。C2D projection通过估计最优投影矩阵,将近距麦克风输入转换为与远距录音对齐的干净参考信号,同时实现降噪。该投影可通过参数化多通道维纳滤波器(PMWF)变体有效实现。实验表明,在使用GSS增强输出作为辅助输入的条件下,基于C2D投影数据训练的神经网络在具有挑战性的CHiME6聚餐场景语音识别任务中表现优于当前最优的引导源分离(GSS)方法。
原文摘要 · Abstract (English)
Training neural networks (NNs) for speech enhancement (SE) in distant speech-capturing scenarios requires paired distorted and clean reference speech signals. While such data are often generated through simulation, the mismatch between simulated and real recordings significantly limits SE accuracy. To address this issue, we propose Close-to-Distant microphone Projection (C2D projection), a method that generates paired data from real recordings captured by close and distant microphones. C2D projection estimates an optimal projection matrix that transforms close-microphone inputs into clean reference signals aligned with distant-microphone recordings, while simultaneously performing denoising. We show this projection can be effectively realized using a variant of the Parametric Multichannel Wiener Filter (PMWF). Experimental results demonstrate that an NN trained with C2D-projected data outperforms the state-of-the-art Guided Source Separation (GSS) on the challenging CHiME6 dinner party ASR task under oracle diarization, when using the enhanced output from GSS as an auxiliary input to the NN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。