arXiv:2411.06307cs.SDeess.AS2024-11NeurIPS被引 30

用体积渲染技术生成更真实的音频响应,提升虚拟现实听觉沉浸感。

Acoustic Volume Rendering for Neural Impulse Response Fields

  • 将体积渲染拓展到频域,建模声音在空间中的传播路径。
  • 在新视角下合成的冲激响应性能超越现有方法,误差更低。
  • 适合做音效模拟、虚拟现实开发与声学仿真研究者使用。

真实音频合成需精准捕捉声学现象,以实现虚拟与增强现实中的沉浸体验。声音在任意位置的接收依赖于冲激响应(IR)的估计,其描述了声音在场景中沿不同路径传播后到达听者的过程。本文提出声学体积渲染(AVR),首次将体积渲染技术应用于建模冲激响应。由于IR是时间序列信号,传统方法面临挑战,为此我们引入频域体积渲染,并采用球面积分拟合实际测量的IR数据。该方法构建的冲激响应场天然包含波传播原理,在新视角下的合成表现达到当前最优。实验表明,AVR显著优于现有领先方法。此外,我们开发了声学仿真平台AcoustiX,相比现有工具提供更精确、真实的IR模拟。AVR与AcoustiX代码已开源:https://zitonglan.github.io/avr。

原文摘要 · Abstract (English)

Realistic audio synthesis that captures accurate acoustic phenomena is essential for creating immersive experiences in virtual and augmented reality. Synthesizing the sound received at any position relies on the estimation of impulse response (IR), which characterizes how sound propagates in one scene along different paths before arriving at the listener's position. In this paper, we present Acoustic Volume Rendering (AVR), a novel approach that adapts volume rendering techniques to model acoustic impulse responses. While volume rendering has been successful in modeling radiance fields for images and neural scene representations, IRs present unique challenges as time-series signals. To address these challenges, we introduce frequency-domain volume rendering and use spherical integration to fit the IR measurements. Our method constructs an impulse response field that inherently encodes wave propagation principles and achieves state-of-the-art performance in synthesizing impulse responses for novel poses. Experiments show that AVR surpasses current leading methods by a substantial margin. Additionally, we develop an acoustic simulation platform, AcoustiX, which provides more accurate and realistic IR simulations than existing simulators. Code for AVR and AcoustiX are available at https://zitonglan.github.io/avr.

音频生成体积渲染声学仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。