用低内存方法高效还原逼真声学效果,让虚拟场景听感更真实。
Reciprocal Latent Fields for Precomputed Sound Propagation
- 通过可学习的体素嵌入与对称解码器,实现声学参数的高效编码。
- 在复杂场景中保持高保真度,内存占用降低数个数量级。
- 适合需要实时沉浸式音效的虚拟现实、游戏和元宇宙应用。
真实声波传播对虚拟场景的沉浸感至关重要,但基于物理的波场模拟在实时应用中仍计算成本过高。波编码方法通过预计算并压缩场景中各声源-接收点的脉冲响应为标量声学参数,但在大型环境中参数规模庞大。本文提出一种名为互易潜在场(Reciprocal Latent Fields, RLF)的内存高效框架,用于编码与预测这些声学参数。该框架采用可训练的体素网格潜变量,并通过对称函数解码,确保声学互易性。我们研究了多种解码器,发现利用黎曼度量学习能更好复现复杂场景中的声学现象。实验表明,RLF在显著降低内存占用的同时保持高质量还原。此外,类似MUSHRA的主观听觉测试显示,通过RLF生成的声音在感知上与真实模拟结果无差异。
原文摘要 · Abstract (English)
Realistic sound propagation is essential for immersion in a virtual scene, yet physically accurate wave-based simulations remain computationally prohibitive for real-time applications. Wave coding methods address this limitation by precomputing and compressing impulse responses of a given scene into a set of scalar acoustic parameters, which can reach unmanageable sizes in large environments with many source-receiver pairs. We introduce Reciprocal Latent Fields (RLF), a memory-efficient framework for encoding and predicting these acoustic parameters. The RLF framework employs a volumetric grid of trainable latent embeddings decoded with a symmetric function, ensuring acoustic reciprocity. We study a variety of decoders and show that leveraging Riemannian metric learning leads to a better reproduction of acoustic phenomena in complex scenes. Experimental validation demonstrates that RLF maintains replication quality while reducing the memory footprint by several orders of magnitude. Furthermore, a MUSHRA-like subjective listening test indicates that sound rendered via RLF is perceptually indistinguishable from ground-truth simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。