arXiv:2410.22805cs.SDcs.AI2024-10中稿 · APSIPA2024被引 1

提出实时自适应神经波束成形,提升复杂环境下的语音增强效果。

Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising

  • 用联合去混响与去噪的WPD波束成形,让神经网络可在线调整。
  • 在多说话人、不同混响时间与信噪比下,显著降低识别错误率。
  • 适合部署于实际场景的实时语音识别系统,如智能音箱和会议设备。

本文针对真实环境中实时自动语音识别(ASR)的语音增强问题,提出一种可运行时自适应的神经波束成形方法。标准方法使用深度神经网络(DNN)从含噪回声混合谱图中估计干净干声掩码,并生成波束成形增强滤波器,但在条件不匹配时性能急剧下降。为此,本文引入运行时自适应机制:虽无真实语音谱图可用,但可通过盲去混响与分离方法(如加权预测误差WPE、快速多通道非负矩阵分解FastMNMF)生成伪真值数据。此前工作采用异步级联的WPE与最小方差无失真响应(MVDR)波束成形,由块在线的FastMNMF微调。本文则提出统一的加权功率最小化无失真响应(WPD)波束成形,其联合去混响与去噪滤波器由DNN估计,实现更紧密的集成与在线可调性。通过在不同说话人数、混响时间及信噪比(SNR)条件下评估,验证了该方法在多种复杂场景下的鲁棒性与有效性。

原文摘要 · Abstract (English)

This paper describes speech enhancement for realtime automatic speech recognition (ASR) in real environments. A standard approach to this task is to use neural beamforming that can work efficiently in an online manner. It estimates the masks of clean dry speech from a noisy echoic mixture spectrogram with a deep neural network (DNN) and then computes a enhancement filter used for beamforming. The performance of such a supervised approach, however, is drastically degraded under mismatched conditions. This calls for run-time adaptation of the DNN. Although the ground-truth speech spectrogram required for adaptation is not available at run time, blind dereverberation and separation methods such as weighted prediction error (WPE) and fast multichannel nonnegative matrix factorization (FastMNMF) can be used for generating pseudo groundtruth data from a mixture. Based on this idea, a prior work proposed a dual-process system based on a cascade of WPE and minimum variance distortionless response (MVDR) beamforming asynchronously fine-tuned by block-online FastMNMF. To integrate the dereverberation capability into neural beamforming and make it fine-tunable at run time, we propose to use weighted power minimization distortionless response (WPD) beamforming, a unified version of WPE and minimum power distortionless response (MPDR), whose joint dereverberation and denoising filter is estimated using a DNN. We evaluated the impact of run-time adaptation under various conditions with different numbers of speakers, reverberation times, and signal-to-noise ratios (SNRs).

语音增强神经波束成形实时处理自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。