arXiv:2604.05545eess.AS2026-04中稿 · ICASSP 2026

用深度学习实时生成空间混响,提升虚拟现实听觉真实感

Multimodal Deep Learning Method for Real-Time Spatial Room Impulse Response Computing

  • 输入场景信息与低阶反射波形,联合建模空间声学特性
  • 相比传统方法,生成速度更快且更适配个性化耳机效果
  • 新构建多样化数据集,支持复杂场景的高保真还原

我们提出一种用于虚拟现实听觉再现的多模态深度学习模型,可实时生成空间房间冲激响应(SRIR),以重建特定场景的听觉感知。将SRIR作为输出能降低计算复杂度,并便于与个性化头相关传递函数结合。模型输入包含场景信息和波形,其中波形对应于低阶反射(LoR)。LoR可通过几何声学(GA)高效计算,但对深度学习模型而言难以准确预测。首先利用场景几何、声学属性、声源与听者坐标实时通过GA计算出LoR,再将LoR及这些特征一同输入模型。为此构建了一个包含多个场景及其对应SRIR的新数据集,具有更高多样性。实验结果表明,所提模型性能显著优于现有方法。

原文摘要 · Abstract (English)

We propose a multimodal deep learning model for VR auralization that generates spatial room impulse responses (SRIRs) in real time to reconstruct scene-specific auditory perception. Employing SRIRs as the output reduces computational complexity and facilitates integration with personalized head-related transfer functions. The model takes two modalities as input: scene information and waveforms, where the waveform corresponds to the low-order reflections (LoR). LoR can be efficiently computed using geometrical acoustics (GA) but remains difficult for deep learning models to predict accurately. Scene geometry, acoustic properties, source coordinates, and listener coordinates are first used to compute LoR in real time via GA, and both LoR and these features are subsequently provided as inputs to the model. A new dataset was constructed, consisting of multiple scenes and their corresponding SRIRs. The dataset exhibits greater diversity. Experimental results demonstrate the superior performance of the proposed model.

虚拟现实声学建模深度学习实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。