arXiv:2507.15229eess.AS2025-07

用波束成形混合信号作弱监督,提升语音增强与降噪识别效果

Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR

  • 用高信噪比的波束成形混合信号作为弱标签训练增强模型
  • 在CHiME-4真实数据集上显著提升语音增强和语音识别性能
  • 适合需要真实场景泛化能力的语音增强与语音识别研究者

在多通道语音增强与鲁棒自动语音识别中,波束成形通常能提高目标说话人的信噪比(SNR),并生成对目标语音失真小的可靠增强结果。基于此观察,本文提出利用波束成形混合信号(其目标说话人信噪比高于原始混合信号)作为弱监督信号,训练深度神经网络以增强输入混合信号。该方法可使用真实录制的混合信号与其波束成形版本配对进行训练,相比仅在模拟混合信号上训练的模型,有望实现更好的真实混合信号泛化能力。在真实录制的CHiME-4数据集上的评估结果验证了该方法的有效性。

原文摘要 · Abstract (English)

In multi-channel speech enhancement and robust automatic speech recognition (ASR), beamforming can typically improve the signal-to-noise ratio (SNR) of the target speaker and produce reliable enhancement with little distortion to target speech. With this observation, we propose to leverage beamformed mixture, which has a higher SNR of the target speaker than the input mixture, as a weak supervision to train deep neural networks (DNNs) to enhance the input mixture. This way, we can train enhancement models using pairs of real-recorded mixture and its beamformed mixture, and potentially realize better generalization to real mixtures, compared with only training the models on simulated mixtures, which usually mismatch real mixtures. Evaluation results on the real-recorded CHiME-4 dataset show the effectiveness of the proposed algorithm.

语音增强波束成形弱监督语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。