用神经多极子模拟声场,快速生成任意位置的混响响应。
Neural acoustic multipole splatting for room impulse response synthesis
- 用神经网络学习多极子位置与声源特性,构建声场模型。
- 在真实和合成数据上优于现有方法,20%极子数达单极子性能。
- 适合需要高效声学建模的虚拟现实与空间音频应用。
在任意接收位置预测房间冲激响应(RIR)对空间音频渲染等应用至关重要。本文提出神经声学多极子点阵(NAMS),通过学习神经声学多极子的位置,并利用神经网络预测其发射信号与指向性,合成未见位置的RIR。将声场表示为多极子组合,在满足亥姆霍兹方程等物理约束的同时,具备表达复杂声学场景的灵活性。我们还引入一种剪枝策略,从密集点阵出发,在训练中逐步剔除冗余多极子。在真实与合成数据上的实验表明,该方法在多数指标上超越先前方法,且推理速度迅速。消融实验显示,采用剪枝的多极子点阵仅需20%极子数,即可达到单极子模型的性能。
原文摘要 · Abstract (English)
Room Impulse Response (RIR) prediction at arbitrary receiver positions is essential for practical applications such as spatial audio rendering. We propose Neural Acoustic Multipole Splatting (NAMS), which synthesizes RIRs at unseen receiver positions by learning the positions of neural acoustic multipoles and predicting their emitted signals and directivities using a neural network. Representing sound fields through a combination of multipoles offers sufficient flexibility to express complex acoustic scenes while adhering to physical constraints such as the Helmholtz equation. We also introduce a pruning strategy that starts from a dense splatting of neural acoustic multipoles and progressively eliminates redundant ones during training. Experiments conducted on both real and synthetic datasets indicate that the proposed method surpasses previous approaches on most metrics while maintaining rapid inference. Ablation studies reveal that multipole splatting with pruning achieves better performance than the monopole model with just 20% of the poles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。