arXiv:2608.24558eess.ASeess.SP2026-08

用生成模型补偿麦克风阵列差异,任意阵列都能一键生成高质量3D音频。

Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling

论文配图:Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling
图 1 · 摘自论文原文
  • 将物理录音模型嵌入推理过程,自适应修正阵列带来的失真
  • 在多种真实与模拟阵列上测试,音质和空间保真度均超越传统方法
  • 无需重新训练,可直接用于任意阵列拓扑,适合音频制作与虚拟现实场景

空间音频通过还原三维声场增强用户沉浸感,其中全向声学(Ambisonics)是一种广泛应用的表示方式。尽管理论上与录音设备无关,实际麦克风阵列仍会引入硬件相关的编码畸变。现有数据驱动方法灵活性不足,通常仅适用于固定阵列几何结构。为此,我们提出ADEPS,一种生成式框架,将物理采集模型显式嵌入推理过程。借助该设计,ADEPS能有效补偿阵列特异性失真,并实现对任意阵列拓扑的零样本编码。我们在无监督条件下仅基于目标Ambisonic表示训练底层生成先验。在多种模拟与真实麦克风阵列上的广泛评估表明,ADEPS在空间保真度与频谱质量方面持续优于传统线性与参数化基线。

原文摘要 · Abstract (English)

Spatial audio enhances user immersion by reproducing 3D sound fields, with Ambisonics being a widely adopted representation. While Ambisonics is theoretically independent of the recording setup, practical microphone arrays introduce hardware-dependent encoding artifacts. Moreover, existing data-driven solutions lack flexibility, as they are typically restricted to fixed array geometries. To overcome these limitations, we propose ADEPS, a generative framework that explicitly embeds the physical acquisition model into the inference process. By leveraging this formulation, ADEPS effectively compensates for array-specific distortions while enabling zero-shot encoding across arbitrary array topologies. We train the underlying generative prior in an unsupervised manner solely on target Ambisonic representations. Extensive evaluations across diverse simulated and real microphone arrays demonstrate that ADEPS consistently outperforms both traditional linear and parametric baselines in spatial fidelity and spectral quality.

空间音频生成模型麦克风阵列3D声场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。