用复数网络提升双耳语音增强,更好保留声音方向感。
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
- 用复数卷积循环网络建模双耳信号,分左右耳估计复数谱掩码。
- 在单说话人、各类型噪声下,显著提升语音可懂度并降噪。
- 特别适合助听器、虚拟现实等需保空间信息的场景。
从助听器到增强与虚拟现实设备,双耳语音增强技术已成为提升语音可懂度和听觉舒适度的主流方法。本文提出一种基于编码器-解码器架构与复数长短期记忆循环模块的端到端双耳语音增强方法。通过引入关注空间信息保持的损失函数,网络在时频域分别估计左、右耳通道的复数比掩码。实验表明,在单目标说话人与各类型各向同性噪声环境下,相比现有基线算法,该方法显著提升了语音可懂度,有效降低噪声,同时完整保留双耳信号的空间特性。
原文摘要 · Abstract (English)
From hearing aids to augmented and virtual reality devices, binaural speech enhancement algorithms have been established as state-of-the-art techniques to improve speech intelligibility and listening comfort. In this paper, we present an end-to-end binaural speech enhancement method using a complex recurrent convolutional network with an encoder-decoder architecture and a complex LSTM recurrent block placed between the encoder and decoder. A loss function that focuses on the preservation of spatial information in addition to speech intelligibility improvement and noise reduction is introduced. The network estimates individual complex ratio masks for the left and right-ear channels of a binaural hearing device in the time-frequency domain. We show that, compared to other baseline algorithms, the proposed method significantly improves the estimated speech intelligibility and reduces the noise while preserving the spatial information of the binaural signals in acoustic situations with a single target speaker and isotropic noise of various types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。