arXiv:2505.22051eess.AS2025-05中稿 · Interspeech 2025被引 4

用自回归机制提升多通道语音增强效果,让模型参考前帧信息更准地还原语音。

ARiSE: Auto-Regressive Multi-Channel Speech Enhancement

  • 利用前帧语音估计结果作为额外输入,引导当前帧语音重建
  • 在混响噪声环境下显著提升语音清晰度,优于传统端到端模型
  • 提出并行训练方法,解决自回归模型训练慢的痛点,适合语音增强研究者

我们提出 ARiSE,一种用于多通道语音增强的自回归算法。ARiSE 通过引入自回归连接,将前一帧的语音估计结果作为额外输入特征,帮助神经网络更准确地估计当前帧的目标语音。这些额外特征可来自(a)前帧的语音估计值;或(b)基于前帧估计语音计算出的波束成形混合信号。然而,直接以自回归方式训练深度神经网络(DNN)效率极低。为此,我们设计了一种并行训练机制,有效加速训练过程。在噪声与混响环境下的评估结果显示,该方法在语音增强性能上具有显著优势和应用潜力。

原文摘要 · Abstract (English)

We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regressive connections, where the estimated target speech at previous frames is leveraged as extra input features to help the DNN estimate the target speech at the current frame. The extra input features can be derived from (a) the estimated target speech in previous frames; and (b) a beamformed mixture with the beamformer computed based on the previous estimated target speech. On the other hand, naively training the DNN in an auto-regressive manner is very slow. To deal with this, we propose a parallel training mechanism to speed up the training. Evaluation results in noisy-reverberant conditions show the effectiveness and potential of the proposed algorithms.

语音增强自回归模型多通道

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。