arXiv:2504.21847cs.CVcs.SD2025-04ICCV被引 7

用多视角视觉信息提升声学渲染效率与精度

Differentiable Room Acoustic Rendering with Multi-View Vision Priors

  • 融合多视角图像与声线追踪,实现可微分的物理声学建模
  • 在真实场景中性能超越多个基线方法,训练数据仅需1/10仍达相当水平
  • 适合需要高保真音频渲染的虚拟现实与数字孪生应用

沉浸式声学体验在构建真实虚拟环境中的重要性不亚于视觉效果。然而,现有房间冲击响应估计方法要么依赖数据密集型学习模型,要么计算成本高昂。本文提出音视频可微分声学渲染(AV-DAR)框架,利用多视角图像提取视觉线索,并结合声线追踪实现基于物理的声学渲染。在两个数据集共六个真实场景上的实验表明,该多模态、基于物理的方法高效、可解释且准确,显著优于多个先前方法。特别地,在Real Acoustic Field数据集上,AV-DAR在相同训练规模下相比基线模型相对提升16.6%至50.9%,性能接近训练数据量为10倍的模型。

原文摘要 · Abstract (English)

An immersive acoustic experience enabled by spatial audio is just as crucial as the visual aspect in creating realistic virtual environments. However, existing methods for room impulse response estimation rely either on data-demanding learning-based models or computationally expensive physics-based modeling. In this work, we introduce Audio-Visual Differentiable Room Acoustic Rendering (AV-DAR), a framework that leverages visual cues extracted from multi-view images and acoustic beam tracing for physics-based room acoustic rendering. Experiments across six real-world environments from two datasets demonstrate that our multimodal, physics-based approach is efficient, interpretable, and accurate, significantly outperforming a series of prior methods. Notably, on the Real Acoustic Field dataset, AV-DAR achieves comparable performance to models trained on 10 times more data while delivering relative gains ranging from 16.6% to 50.9% when trained at the same scale.

声学渲染多模态可微分虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。