arXiv:2506.06190cs.SDcs.GR2025-06被引 1

用神经网络实时模拟动态场景中的声音变化。

NAT: Neural Acoustic Transfer for Interactive Scenes in Real Time

  • 用隐式神经表示编码声学传递及变化,支持实时预测。
  • 训练数据生成速度提升,30秒音频处理仅需数毫秒。
  • 适合虚拟现实、增强现实等需要实时音效的场景。

以往声学传递方法依赖大量预计算和存储数据以实现实时交互与听觉反馈,但在物体位置、材质、尺寸动态变化的复杂场景中表现不佳,导致声学传递分布持续波动,难以用基础数据结构高效表示与渲染。为此,我们提出神经声学传递(NAT),采用隐式神经表示编码预计算的声学传递及其变化,实现对不同条件下声场的实时预测。为高效生成训练数据,我们开发了一种基于快速蒙特卡洛的边界元法(BEM)近似方法,适用于具有光滑诺伊曼条件的一般场景;同时实现了高精度标准BEM的GPU加速版本。这些方法提供了必要的训练数据,使神经网络能准确建模声辐射空间。通过在多种声学传递场景中的全面验证与对比,我们的方法在数值精度和运行效率上表现优异,30秒音频处理时间控制在数毫秒内。该方法可高效精准地建模动态环境中的声音行为,适用于虚拟现实、增强现实及高级音频制作等交互应用。

原文摘要 · Abstract (English)

Previous acoustic transfer methods rely on extensive precomputation and storage of data to enable real-time interaction and auditory feedback. However, these methods struggle with complex scenes, especially when dynamic changes in object position, material, and size significantly alter sound effects. These continuous variations lead to fluctuating acoustic transfer distributions, making it challenging to represent with basic data structures and render efficiently in real time. To address this challenge, we present Neural Acoustic Transfer, a novel approach that utilizes an implicit neural representation to encode precomputed acoustic transfer and its variations, allowing for real-time prediction of sound fields under varying conditions. To efficiently generate the training data required for the neural acoustic field, we developed a fast Monte-Carlo-based boundary element method (BEM) approximation for general scenarios with smooth Neumann conditions. Additionally, we implemented a GPU-accelerated version of standard BEM for scenarios requiring higher precision. These methods provide the necessary training data, enabling our neural network to accurately model the sound radiation space. We demonstrate our method's numerical accuracy and runtime efficiency (within several milliseconds for 30s audio) through comprehensive validation and comparisons in diverse acoustic transfer scenarios. Our approach allows for efficient and accurate modeling of sound behavior in dynamically changing environments, which can benefit a wide range of interactive applications such as virtual reality, augmented reality, and advanced audio production.

声学模拟神经渲染实时交互虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。