为分散麦克风设计多通道掩码,提升语音增强效果。
A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones
- 每麦克风独立加掩码,捕捉空间差异性
- 信号差异大时性能优于单掩码,紧凑阵列无损失
- 适合远场、动态声学场景的语音增强系统
基于掩码的波束成形是一种不依赖几何结构的语音增强方法,通常对所有麦克风使用单一掩码来估计协方差矩阵。然而,对于空间分布式的麦克风,这种策略可能次优,因为各麦克风信号特性差异显著。为此,本文将掩码波束成形扩展为多通道形式:在协方差估计前,对每个麦克风分别施加独立掩码,以更好捕捉空间多样性。针对时变声学场景(由频谱-时序非平稳性引起),采用滑动窗口的帧因果在线实现。在模拟紧凑阵列与分布式麦克风上的实验表明,当麦克风信号差异显著时,多通道掩码优于单掩码;而在紧凑阵列中性能相当。进一步通过对比理想比率掩码与盲深度神经网络掩码估计,验证了该方法的鲁棒性。
原文摘要 · Abstract (English)
Mask-based beamforming is a popular geometry-agnostic approach for speech enhancement, typically applying a single mask across all microphones to estimate the required covariance matrices. While effective for compact arrays, this strategy may be suboptimal for spatially distributed microphones, where signal characteristics may vary strongly across microphones. To effectively capture the spatial diversity across microphones, we extend the mask-based beamformer to a multi-channel formulation, where each microphone is pre-filtered by a separate mask before covariance estimation. To address time-varying acoustic scenes, caused by spectro-temporal nonstationarity, we adopt a frame-causal online implementation with a sliding window. Experiments with simulated compact arrays and distributed microphones show that multi-channel masking yields a benefit over using a single mask when microphone signals differ substantially, while retaining similar performance in compact arrays. We further demonstrate the robustness of the multi-channel masking approach by comparing oracle ideal ratio masks to blind DNN-based mask estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。