arXiv:2502.07205eess.AScs.AI2025-02被引 3

用神经语音先验联合去混响与声学环境建模,提升语音识别效果。

VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification

  • 基于变分贝叶斯框架,用深度网络估计语音先验分布。
  • 在单通道下实现去混响与混响脉冲响应盲辨识,性能达顶尖水平。
  • 适合语音识别系统优化,尤其对混响严重场景有显著提升。

混响语音包含直达声和房间脉冲响应(RIR)的关键信息。本文提出一种基于变分贝叶斯推断(VBI)的神经语音先验框架(VINP),用于联合去混响与盲RIR辨识。在时频域基于卷积传输函数(CTF)近似构建概率信号模型,首次利用任意判别性深度神经网络(DNN)估计无混响语音的先验分布。结合混响语音与先验信息,VINP分别得到无混响语音谱和CTF滤波器的最大后验(MAP)与最大似然(ML)估计,并通过简单变换还原无混响语音波形与RIR。该方法显著提升自动语音识别(ASR)性能,优于多数单通道深度学习去混响方法。单通道实验表明,VINP在平均意见分(MOS)和词错误率(WER)上达到当前最优(SOTA)。在盲RIR辨识方面,其在60 dB衰减时间(RT60)估计上达SOTA,直接-混响比(DRR)估计表现优异。代码与音频样本已公开。

原文摘要 · Abstract (English)

Reverberant speech, denoting the speech signal degraded by reverberation, contains crucial knowledge of both anechoic source speech and room impulse response (RIR). This work proposes a variational Bayesian inference (VBI) framework with neural speech prior (VINP) for joint speech dereverberation and blind RIR identification. In VINP, a probabilistic signal model is constructed in the time-frequency (T-F) domain based on convolution transfer function (CTF) approximation. For the first time, we propose using an arbitrary discriminative dereverberation deep neural network (DNN) to estimate the prior distribution of anechoic speech within a probabilistic model. By integrating both reverberant speech and the anechoic speech prior, VINP yields the maximum a posteriori (MAP) and maximum likelihood (ML) estimations of the anechoic speech spectrum and CTF filter, respectively. After simple transformations, the waveforms of anechoic speech and RIR are estimated. VINP is effective for automatic speech recognition (ASR) systems, which sets it apart from most deep learning (DL)-based single-channel dereverberation approaches. Experiments on single-channel speech dereverberation demonstrate that VINP attains state-of-the-art (SOTA) performance in mean opinion score (MOS) and word error rate (WER). For blind RIR identification, experiments demonstrate that VINP achieves SOTA performance in estimating reverberation time at 60 dB (RT60) and advanced performance in direct-to-reverberation ratio (DRR) estimation. Codes and audio samples are available online.

语音去混响贝叶斯推理深度学习语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。