无需声源先验,通过视觉-声音绑定生成新视角的环境音。
SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
- 用全景图像和深度数据学习视觉-声学关联,实现跨视角声音合成。
- 在真实场景中比现有方法显著提升音质与空间一致性。
- 适合做虚拟现实、影视音效生成的开发者或研究者。
我们提出 SoundVista,一种从任意新视角生成场景环境音的方法。给定稀疏分布麦克风采集的原始录音,SoundVista 可合成该场景在未见目标视角下的声音。该方法学习分布式麦克风信号与目标视角信号之间的隐式声学传递函数,仅需少量已知录音即可完成建模。与现有工作不同,本方法无需声源细节的约束或先验知识。同时具备对多样房间布局、参考麦克风配置及未见环境的高效适应能力。为此,我们引入视觉-声学绑定模块,从全景RGB与深度数据中学习与局部声学特性关联的视觉嵌入。首先利用这些嵌入优化任意场景中的参考麦克风布置;在合成阶段,通过提取多个参考位置的嵌入,动态计算其对目标视角的贡献权重。我们在公开数据集与真实场景中进行基准测试,结果表明该方法显著优于现有技术。
原文摘要 · Abstract (English)
We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。