arXiv:2503.12806cs.MMcs.CV2025-03被引 3

利用表面法向增强几何感知,提升新视角音频合成真实度

AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis

  • 结合表面法向与结构细节建模声波反射路径
  • 在RWAVS和SoundSpace数据集上超越现有方法
  • 适合音频视觉融合与虚拟现实场景开发者

准确建模复杂真实环境中的声音传播对新视角音频合成(NVAS)至关重要。尽管先前研究借助视觉感知估计空间声学特性,但尚未充分探索3D表示中表面法向与结构细节在声学建模中的联合应用。由于它们直接影响声波反射与传播,应联合建模以实现精确的空间声学。本文提出一种表面增强的几何感知新视角音频合成方法,利用3D高斯溅射(3DGS)框架获取的图像、深度图、表面法向和点云等几何先验信息。设计基于双交叉注意力的Transformer,将几何约束融入频率查询,以理解发射源周围的环境。同时构建基于ConvNeXt的频谱特征处理网络——频谱精炼网络(SRN),用于生成逼真的双耳音频。在RWAVS和SoundSpace数据集上的实验表明,该方法显著优于现有方法,验证了其有效性。

原文摘要 · Abstract (English)

Accurately modeling sound propagation with complex real-world environments is essential for Novel View Acoustic Synthesis (NVAS). While previous studies have leveraged visual perception to estimate spatial acoustics, the combined use of surface normal and structural details from 3D representations in acoustic modeling has been underexplored. Given their direct impact on sound wave reflections and propagation, surface normals should be jointly modeled with structural details to achieve accurate spatial acoustics. In this paper, we propose a surface-enhanced geometry-aware approach for NVAS to improve spatial acoustic modeling. To achieve this, we exploit geometric priors, such as image, depth map, surface normals, and point clouds obtained using a 3D Gaussian Splatting (3DGS) based framework. We introduce a dual cross-attention-based transformer integrating geometrical constraints into frequency query to understand the surroundings of the emitter. Additionally, we design a ConvNeXt-based spectral features processing network called Spectral Refinement Network (SRN) to synthesize realistic binaural audio. Experimental results on the RWAVS and SoundSpace datasets highlight the necessity of our approach, as it surpasses existing methods in novel view acoustic synthesis.

音频合成几何感知3DGS双耳音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。