用U-Net网络预测自适应去混叠滤波器,提升麦克风阵列音视频采集的空间精度。
Deep learning based spatial aliasing reduction in beamforming for audio capture
- 基于U-Net构建信号相关去混叠滤波器,学习跨通道依赖关系。
- 在立体声与一阶Ambisonics场景中,客观与主观表现显著优于传统波束成形。
- 适用于多麦克风系统,尤其适合空间混叠严重的音频捕获场景。
空间混叠影响间隔麦克风阵列,在特定频率以上引发方向模糊,降低波束成形的空间与频谱准确性。针对传统信号处理的局限性及深度学习在去混叠方面的稀缺,本文提出一种新方法:利用U-Net架构预测信号依赖的去混叠滤波器,用于改进传统波束成形的空间捕获。考虑两种多通道滤波器:一种独立处理各通道,另一种建模通道间依赖。在立体声与一阶Ambisonics两种常见空间捕获场景中评估,结果表明该方法在客观指标和听觉感知上均实现显著提升。本工作展示了深度学习在减少波束成形混叠方面的潜力,可有效改善多麦克风系统的音频质量。
原文摘要 · Abstract (English)
Spatial aliasing affects spaced microphone arrays, causing directional ambiguity above certain frequencies, degrading spatial and spectral accuracy of beamformers. Given the limitations of conventional signal processing and the scarcity of deep learning approaches to spatial aliasing mitigation, we propose a novel approach using a U-Net architecture to predict a signal-dependent de-aliasing filter, which reduces aliasing in conventional beamforming for spatial capture. Two types of multichannel filters are considered, one which treats the channels independently and a second one that models cross-channel dependencies. The proposed approach is evaluated in two common spatial capture scenarios: stereo and first-order Ambisonics. The results indicate a very significant improvement, both objective and perceptual, with respect to conventional beamforming. This work shows the potential of deep learning to reduce aliasing in beamforming, leading to improvements in multi-microphone setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。