轻量级模型提升双耳语音增强,兼顾降噪与空间线索保真
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
- 基于低频滤波与通道间声学传递函数估计,实现高效降噪
- 在固定说话人条件下达到顶尖降噪性能,计算成本显著降低
- 适合实时听力设备应用,尤其关注空间听觉自然性的场景
双耳语音增强旨在提升助听设备接收的噪声信号的语音质量与可懂度,同时保留目标声源的空间线索以实现自然聆听。现有方法常在降噪能力与空间线索保真之间存在权衡,且在复杂声学场景下计算开销高。本文提出一种基于学习的轻量级双耳复数卷积网络(LBCCN),通过滤除低频分量实现优异降噪效果,同时显式估计通道间相对声学传递函数,保障空间线索准确性和语音清晰度。实验表明,在固定说话人条件下,该模型降噪性能媲美当前最优方法,但计算成本大幅降低,并具备一定空间线索保持能力。代码与音频示例已公开于 https://github.com/jywanng/LBCCN。
原文摘要 · Abstract (English)
Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer from the compromise between noise reduction (NR) capacity and spatial cues preservation (SCP) accuracy and a high computational demand in complex acoustic scenes. In this work, we present a learning-based lightweight binaural complex convolutional network (LBCCN), which excels in NR by filtering low-frequency bands and keeping the rest. Additionally, our approach explicitly incorporates the estimation of interchannel relative acoustic transfer function to ensure the spatial cues fidelity and speech clarity. Results show that the proposed LBCCN can achieve a comparable NR performance to state-of-the-art methods under fixed-speaker conditions, but with a much lower computational cost and a certain degree of SCP capability. The reproducible code and audio examples are available at https://github.com/jywanng/LBCCN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。