用速度势函数建模声场,让空间音频信号自动满足物理规律。
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
- 用单通道速度势函数替代直接预测四通道音信号
- 通过求导生成四通道信号,保证任意时空点均符合线性动量方程
- 在房间混响重建中表现更优,适合追求物理一致性的音频应用
一阶全向声学(FOA)是基于球谐分解的标准空间音频格式,其零阶与一阶分量分别表征声压与粒子速度。近年来,物理信息神经网络被用于对FOA信号进行空间插值,通过软约束项将物理原理(如线性动量方程)融入网络输出。本文提出新方法:不再直接预测FOA信号,而是学习一个标量函数——速度势。该函数的偏导数可直接生成四通道FOA信号,且在任意时间与麦克风位置上天然满足线性动量方程。实验结果表明,该框架在房间冲激响应重建任务中具有显著有效性。
原文摘要 · Abstract (English)
First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and particle velocity, respectively. Recently, physics-informed neural networks have been applied to the spatial interpolation of FOA signals, regularizing the network outputs based on soft penalty terms derived from physical principles, e.g., the linearized momentum equation. In this paper, we reformulate the task so that the predicted FOA signal automatically satisfies the linearized momentum equation. Our network approximates a scalar function called velocity potential, rather than the FOA signal itself. Then, the FOA signal can be readily recovered through the partial derivatives of the velocity potential with respect to the network inputs (i.e., time and microphone position) according to physics of sound propagation. By deriving the four channels of FOA from the single-channel velocity potential, the reconstructed signal follows the physical principle at any time and position by construction. Experimental results on room impulse response reconstruction confirm the effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。