arXiv:2601.19712cs.SDcs.MM2026-01中稿 · , Project page: ht…被引 1

让声音更真实:结合视觉语言与物理模型生成沉浸式立体声

Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling

  • 用多视角图像重建3D声学环境,捕捉房间几何结构
  • 融合视觉语言模型提取材质/布局的物理特性,提升声波反射吸收模拟
  • 适合做虚拟现实、影视音效生成的研究者和开发者

空间音频对沉浸式体验至关重要,但新视角声学合成(NVAS)仍面临反射、衍射和材料吸声等复杂物理现象的挑战。现有基于单视角或全景输入的方法虽提升空间保真度,却难以捕捉全局几何与语义线索,如物体布局和材料属性。为此,我们提出首个物理感知的NVAS框架Phys-NVAS,融合空间几何建模与视觉-语言语义先验。通过多视角图像与深度图重建全局3D声学环境,估计房间大小与形状,增强声音传播的空间感知。同时,利用视觉-语言模型提取物体、布局与材料的物理感知先验,捕捉几何之外的吸收与反射特性。声学特征融合适配器将这些线索统一为物理感知表示,用于双耳音频生成。在RWAVS数据集上的实验表明,Phys-NVAS生成的双耳音频更具真实感与物理一致性。

原文摘要 · Abstract (English)

Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on single-view or panoramic inputs improve spatial fidelity but fail to capture global geometry and semantic cues such as object layout and material properties. To address this, we propose Phys-NVAS, the first physics-aware NVAS framework that integrates spatial geometry modeling with vision-language semantic priors. A global 3D acoustic environment is reconstructed from multi-view images and depth maps to estimate room size and shape, enhancing spatial awareness of sound propagation. Meanwhile, a vision-language model extracts physics-aware priors of objects, layouts, and materials, capturing absorption and reflection beyond geometry. An acoustic feature fusion adapter unifies these cues into a physics-aware representation for binaural generation. Experiments on RWAVS demonstrate that Phys-NVAS yields binaural audio with improved realism and physical consistency.

声学合成3D环境建模视觉语言模型双耳音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。