用神经网络实现任意麦克风阵列的声场编码,提升真实设备还原精度。
Beyond Omnidirectional: Neural Ambisonics Encoding for Arbitrary Microphone Directivity Patterns using Cross-Attention
- 通过交叉注意力融合音频与方向响应特征,生成与阵列无关的空间音频表示
- 在混响环境中对多声源模拟测试,性能超越传统方法和现有深度学习方案
- 使用阵列传输函数替代几何信息作为输入,显著提升真实设备建模准确率
我们提出一种深度神经网络方法,将麦克风阵列信号编码为Ambisonics,可泛化至固定麦克风数量但位置各异、具有频率依赖性方向特性的任意阵列配置。不同于以往仅依赖阵列几何结构作为元数据的方法,本方法采用方向性阵列传输函数(array transfer functions),能更准确刻画真实阵列特性。所提架构分别对音频信号和方向响应进行编码,并通过交叉注意力机制融合,生成与阵列无关的空间音频表示。我们在两种模拟场景下评估该方法:带有复杂身体散射的移动设备环境,以及自由场条件,均在混响环境中测试了不同数量声源的情况。实验表明,该方法在性能上优于基于数字信号处理的传统方法及现有深度神经网络解决方案。此外,使用阵列传输函数而非几何结构作为元数据输入,显著提升了真实阵列的建模准确性。
原文摘要 · Abstract (English)
We present a deep neural network approach for encoding microphone array signals into Ambisonics that generalizes to arbitrary microphone array configurations with fixed microphone count but varying locations and frequency-dependent directional characteristics. Unlike previous methods that rely only on array geometry as metadata, our approach uses directional array transfer functions, enabling accurate characterization of real-world arrays. The proposed architecture employs separate encoders for audio and directional responses, combining them through cross-attention mechanisms to generate array-independent spatial audio representations. We evaluate the method on simulated data in two settings: a mobile phone with complex body scattering, and a free-field condition, both with varying numbers of sound sources in reverberant environments. Evaluations demonstrate that our approach outperforms both conventional digital signal processing-based methods and existing deep neural network solutions. Furthermore, using array transfer functions instead of geometry as metadata input improves accuracy on realistic arrays.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。