arXiv:2608.16240eess.AScs.SD2026-08

用物理引导的扩散模型,让稀疏麦克风阵列也能稳定生成高质量空间音频。

Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion

论文配图:Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion
图 1 · 摘自论文原文
  • 基于几何感知的投影方法,将不同布局麦克风信号统一映射到标准球谐表示。
  • 在模拟与真实数据上,对一阶和二阶空间音频编码均显著提升保真度与方向一致性。
  • 适用于可穿戴设备,对未知阵列布局和边界条件有强鲁棒性,适合嵌入式场景。

Ambisonics 提供紧凑的基于场景的空间音频表示,但高阶编码对可穿戴设备和嵌入式硬件构成挑战。其麦克风阵列常稀疏、不规则且受设备边界约束,导致球谐(SH)域编码病态:逆滤波放大噪声,而确定性神经编码器可能过拟合阵列特异性响应或模糊高阶成分。本文提出 DiffM2A,一种面向变拓扑稀疏阵列的几何自适应条件扩散框架,实现鲁棒的 Ambisonic 编码。其前端几何自适应球谐投影(GASHP)构建边界感知的 SH 方向函数,并采用能量归一化模态投影,无需显式伪逆即可将阵列相关观测映射至统一模态表示。双分支清晰化扩散模型则以原始麦克风频谱和 GASHP 特征为条件,估计复杂 Ambisonic 系数。声强与旋转等变损失进一步提升通道间相位一致性和各子空间结构行为。在模拟混响环境与真实 LOCATA 数据上的评估显示,DiffM2A 在信号保真度、频谱精度、空间连贯性及双耳线索保留方面均优于传统与神经基线方法。额外实验表明,该优势在未见过的五麦克风布局以及开阵列与刚性球边界模型不匹配条件下仍保持良好性能。

原文摘要 · Abstract (English)

Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.

空间音频扩散模型麦克风阵列几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。