arXiv:2608.14097eess.AScs.SD2026-08

用扩散模型将不完整麦克风数据转为高阶全向声场,可适配任意设备。

Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model

论文配图:Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model
图 1 · 摘自论文原文
  • 基于扩散模型学习全向声场统计特性,实现任意麦克风阵列的编码。
  • 12阶全向声场重建准确率优于传统与神经基线方法。
  • 适合需跨设备、高保真声学模拟的研究者与开发者。

本文解决从任意且可能不完整或稀疏的麦克风阵列测量中,将房间冲激响应(RIRs)编码为高阶全向声场(HOA)表示的问题。该任务对空间捕获能力有限的麦克风阵列(如不规则或稀疏阵列)而言本质是病态的,经典线性方法无法重建高阶空间细节。我们提出一种基于扩散的生成框架,建模了HOA RIRs的统计特性,从而实现设备无关的编码,即使在训练中未见过的设备也可适用。该方法引入后验采样过程,在估计信号与测量值之间保持一致性的同时,合理重建仅凭有限测量无法观测的空间信息。在模拟数据上的实验表明,本方法优于线性和神经基线,可实现高达12阶的精确HOA RIR估计。包含模拟与实测RIR的双耳渲染听觉测试进一步证实,该方法生成的声场与参考全向声场具有更高的感知相似性。该框架的灵活性与准确性为可扩展声学仿真开辟了新路径。

原文摘要 · Abstract (English)

We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spatial capture capabilities, such as irregular or sparse arrays, as classical linear methods fail to reconstruct high-order spatial detail. We introduce a diffusion-based generative framework that models the statistical properties of HOA RIRs. This enables device-agnostic encoding from arbitrary microphone arrays, potentially unseen during data measurement. Our approach incorporates a posterior sampling procedure that enforces consistency between the estimated signals and the measurements while plausibly reconstructing spatial information that is unobservable from the limited measurements alone. Experiments on simulated data demonstrate that our method outperforms linear and neural baselines, achieving accurate HOA RIR estimation up to 12th order. A listening test with binaural renderings, including both simulated and measured RIRs, further confirms that the proposed method yields higher perceptual similarity to reference Ambisonics RIRs than all baselines. The flexibility and accuracy of the proposed framework opens new possibilities for scalable acoustics simulations.

声学建模扩散模型全向声场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。