让3D场景根据外观生成击打声并定位声音来源。
Visual Acoustic Fields
- 用3D高斯点云融合视觉与声音,建模物体击打声
- 可生成逼真敲击声并精准定位发声位置
- 首个在3D空间中对齐视听信号的数据集
物体被击打时会发出不同声音,人类能根据其外观和材质直观推断声音。受此启发,我们提出视觉声场(Visual Acoustic Fields),利用3D高斯溅射(3DGS)在三维空间中建立击打声与视觉信号的关联。该框架包含两个核心模块:声音生成与声音定位。声音生成模块采用条件扩散模型,以特征增强的3DGS渲染的多尺度特征为输入,生成逼真的击打声;声音定位模块则基于声源信息,查询由特征增强3DGS表示的3D场景,实现击打位置的精确定位。为支持该框架,我们构建了一套新型采集流程,获得场景级的视觉-声音样本对,实现了图像、撞击位置与对应声音的精确对齐。据我们所知,这是首个在3D上下文中连接视觉与声学信号的数据集。在该数据集上的大量实验表明,该方法在生成合理击打声和准确定位声源方面均具有效性。
原文摘要 · Abstract (English)
Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, we propose Visual Acoustic Fields, a framework that bridges hitting sounds and visual signals within a 3D space using 3D Gaussian Splatting (3DGS). Our approach features two key modules: sound generation and sound localization. The sound generation module leverages a conditional diffusion model, which takes multiscale features rendered from a feature-augmented 3DGS to generate realistic hitting sounds. Meanwhile, the sound localization module enables querying the 3D scene, represented by the feature-augmented 3DGS, to localize hitting positions based on the sound sources. To support this framework, we introduce a novel pipeline for collecting scene-level visual-sound sample pairs, achieving alignment between captured images, impact locations, and corresponding sounds. To the best of our knowledge, this is the first dataset to connect visual and acoustic signals in a 3D context. Extensive experiments on our dataset demonstrate the effectiveness of Visual Acoustic Fields in generating plausible impact sounds and accurately localizing impact sources. Our project page is at https://yuelei0428.github.io/projects/Visual-Acoustic-Fields/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。