提升4通道麦克风阵列的声场建模精度,解决低分辨率导致的信息丢失问题。
Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping
- 用轻量卷积网络等方法对稀疏麦克风信号进行超分辨率重建
- 单独训练的轻量模型表现最优,优于联合训练和复杂模型
- 确保重建信号与后续声场建模任务对齐,比模型大小更重要
隐式声场映射(LAM)是一种自监督学习方法,无需标注数据即可从多通道录音生成高分辨率球面声场图,在方向到达基准上达到监督方法水平。然而,当使用稀疏的4通道阵列时,其性能显著下降,因为低分辨率的互谱矩阵所捕捉的空间信息远少于原设计的32通道输入。本文评估了多种超分辨率架构,包括轻量级卷积网络、迭代反投影模型、物理信息网络及生成对抗方法,并研究了将这些上采样器与LAM联合或分阶段训练是否有助于保留LAM依赖的空间结构。结果表明:原始全分辨率LAM仍是最强基线;独立训练的轻量级模型是表现最佳的学习型方法;且上采样器与LAM之间的表征对齐比模型复杂度更为关键。
原文摘要 · Abstract (English)
Latent Acoustic Mapping (LAM) is a self-supervised learning method that generates high-resolution spherical acoustic maps from multichannel recordings without labelled data, matching supervised baselines on direction-of-arrival benchmarks. However, LAM degrades significantly with sparse 4-channel arrays, as the low-resolution cross-spectral matrix captures far less spatial information than the 32-channel inputs LAM was designed for. We benchmark a diverse set of upsampling architectures, spanning lightweight convolutional networks, iterative back-projection models, physics-informed networks, and generative adversarial approaches. We also study whether aligning these upsamplers with LAM by training them jointly or in different stages helps preserve the spatial structure that LAM depends on. Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。