arXiv:2507.06826cs.SDeess.AS2025-07中稿 · WASPAA 2025被引 6

用物理约束提升声场建模精度,实现沉浸式音频生成

Physics-Informed Direction-Aware Neural Acoustic Fields

  • 基于声波传播物理规律,构建FOA四通道间导数关联的先验约束
  • 在真实房间中实验验证,相比普通神经网络提升声场重建准确率
  • 适合做沉浸式音频、虚拟现实音效建模的研究者参考

本文提出一种物理信息神经网络(PINN)来建模一阶全向声学(FOA)房间冲击响应(RIR)。PINN通过结合神经网络的强大拟合能力与声波传播的物理原理,在声场插值方面表现出色。传统方法通常使用全向麦克风测量的声压,利用波动方程或其频域形式——亥姆霍兹方程进行建模。而FOA RIR额外包含空间方向信息,对沉浸式音频生成具有广泛应用价值。本文将PINN框架扩展至建模FOA RIR,基于粒子速度与FOA的(X, Y, Z)通道之间的对应关系,推导出两个物理信息先验。这些先验通过各通道的偏导数关联预测的W通道与其他通道,强制施加四通道间的物理解耦关系。实验表明,所提方法在真实房间环境下显著优于无物理先验的神经网络。

原文摘要 · Abstract (English)

This paper presents a physics-informed neural network (PINN) for modeling first-order Ambisonic (FOA) room impulse responses (RIRs). PINNs have demonstrated promising performance in sound field interpolation by combining the powerful modeling capability of neural networks and the physical principles of sound propagation. In room acoustics, PINNs have typically been trained to represent the sound pressure measured by omnidirectional microphones where the wave equation or its frequency-domain counterpart, i.e., the Helmholtz equation, is leveraged. Meanwhile, FOA RIRs additionally provide spatial characteristics and are useful for immersive audio generation with a wide range of applications. In this paper, we extend the PINN framework to model FOA RIRs. We derive two physics-informed priors for FOA RIRs based on the correspondence between the particle velocity and the (X, Y, Z)-channels of FOA. These priors associate the predicted W-channel and other channels through their partial derivatives and impose the physically feasible relationship on the four channels. Our experiments confirm the effectiveness of the proposed method compared with a neural network without the physics-informed prior.

声场建模物理信息网络沉浸式音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。