用双麦克风阵列实现高噪声下语音增强,动态调整聚焦方向。
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
- 三向引导机制动态确定语音增强范围
- 在极低信噪比下仍保持高语音质量
- 参数少、实时性强,适合设备端部署
在多说话人场景中,利用空间特征对目标语音增强至关重要。然而,在麦克风阵列数量有限的情况下,构建紧凑的多通道语音增强系统仍具挑战性,尤其在极端低信噪比(SNR)条件下。为此,我们提出一种三向引导的空间选择方法,构建了一个灵活框架,使用三个导向向量来指导增强并确定增强范围。具体地,引入因果导向U-Net(CDUNet)模型,以原始多通道语音和期望增强宽度为输入,实现根据目标方向动态调整导向向量,并依据目标与干扰信号之间的角度分离度精细调节增强区域。所提模型仅使用双麦克风阵列,在语音质量和下游任务性能上均表现优异,具备实时运行能力且参数极少,适用于低延迟、设备端流式应用。
原文摘要 · Abstract (English)
In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in extremely low signal-to-noise ratio (SNR) conditions. To tackle this issue, we propose a triple-steering spatial selection method, a flexible framework that uses three steering vectors to guide enhancement and determine the enhancement range. Specifically, we introduce a causal-directed U-Net (CDUNet) model, which takes raw multi-channel speech and the desired enhancement width as inputs. This enables dynamic adjustment of steering vectors based on the target direction and fine-tuning of the enhancement region according to the angular separation between the target and interference signals. Our model with only a dual microphone array, excels in both speech quality and downstream task performance. It operates in real-time with minimal parameters, making it ideal for low-latency, on-device streaming applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。