arXiv:2607.01295eess.AScs.LG2026-07

用深度学习将4麦克风阵列协方差矩阵升维至32通道,提升声成像分辨率。

CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging

论文配图:CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging
图 1 · 摘自论文原文
  • 基于2D卷积与频率动态卷积,从4麦克风输入预测32麦克风协方差矩阵。
  • 最佳模型RMSE达0.432,显著优于随机猜测基线(0.548)。
  • 适合需要高分辨率声成像但受限于硬件的场景,如智能设备、车载系统。

声成像可视化是声学中的核心技术,可实现对声源和声场景的空间分析。然而实际系统中传感器数量有限,促使研究无需增加硬件复杂度即可提升空间分辨率的方法。本文聚焦于通过深度学习技术,将四面体4麦克风阵列虚拟上采样至球形32麦克风阵列,估计各通道协方差矩阵。在真实世界STARSS23数据集上评估了五种神经网络架构,均旨在从4麦克风输入协方差表示预测32麦克风时频协方差矩阵。所提模型基于2D卷积层捕捉协方差矩阵的时空谱结构,并引入频率动态卷积以建模其频率依赖特性。通过均方根误差(RMSE)和延迟-求和波束成形声成像进行评估。定量结果表明,所有模型均优于随机猜测基线(RMSE=0.548),最优模型达到RMSE=0.432。通过波束成形热图可视化,定性分析了模型性能:由4通道输入、32通道真实值及预测输出生成的声图对比显示,协方差上采样显著提升了4通道阵列的有效性能,生成的声图与32通道阵列结果高度一致。

原文摘要 · Abstract (English)

Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes. However, limited sensor availability in practical systems motivate approaches that enhance spatial resolution without increasing the hardware complexity. In this paper, we focus on upsampling virtually a tetrahedral 4-microphone array to a spherical 32-microphone array by estimating the covariance matrices of the channels employing deep learning techniques. Five neural network architectures are investigated for covariance upsampling for acoustic imaging using the real-world STARSS23 dataset. These models are developed to estimate a 32-microphone, time-frequency covariance matrix from a 4-microphone input covariance representation. The proposed architectures are based on 2D convolutional layers to capture the underlying spatial-spectral structure of covariance matrices, and are further enhanced with frequency dynamic convolution to model their frequency-dependent properties. The proposed architectures are evaluated in terms of root mean square error (RMSE) and using delay-and-sum beamforming acoustic imaging. Quantitative results show that all models outperform a random-guess baseline, which yields an RMSE of 0.548, with the best-performing architecture achieving an RMSE of 0.432. We analyze qualitatively the performance of the proposed models through beamforming heatmap visualizations derived from the 4-channel input covariance, the 32-channel ground truth, and the predicted 32-channel covariance matrices. These results demonstrate that covariance upsampling significantly enhances the effective performance of the 4-channel microphone array, producing sound maps that closely resemble those obtained with the 32-channel array.

声成像协方差上采样深度学习麦克风阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。