让3D高斯世界发声,听觉定位让声音随视角移动保持一致。
Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

- 基于多视角检测与高斯集匹配,将声音锚定在3D空间固定位置。
- 在真实与生成的3DGS场景中,声音一致性显著优于单视角基线。
- 无需训练,适用于任意3DGS世界,适合沉浸式应用开发者。
3D高斯泼溅(3DGS)可将捕捉或生成的图像转化为用户可自由探索的逼真3D世界,但这些世界仍无声。现有音频生成方法依赖单张图像或视角,导致声音随观察者移动而失真。本文提出为给定3DGS世界生成空间一致的声音景观,通过听觉定位确定哪些物体发声,并将其锚定于持久的3D位置。我们提出Scene2Sound——一种无需训练的框架,仅从输入世界出发,选取多视角覆盖场景,利用视觉-语言模型识别发声物体,并通过高斯集匹配将多视图检测结果关联为3D实例,该匹配衡量渲染每个检测结果的高斯集重叠度。每个声源接收由标准基于对象的音频引擎实时空间化的声音,支持任意听众姿态。我们还提出两个空间一致性度量:一是验证音频对听众运动响应是否一致;二是检验声源位置是否被排除在外的视角所支持。在精心筛选的生成3DGS世界及真实世界360°拍摄生成的3DGS场景上,Scene2Sound在保持强单视角基线音质的同时,空间一致性显著优于单视角和全景基线,用户研究证实其感知优势。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition on a single image or viewpoint, their sound is tied to that observation and cannot stay consistent while a listener moves. We introduce the task of generating a spatially consistent soundscape for a given 3DGS world through auditory grounding, identifying which objects in the world should emit sound and anchoring each to a persistent 3D position, and present Scene2Sound, a training-free framework built on this grounding. From the input world alone, our pipeline selects viewpoints that jointly cover the scene, identifies sound-emitting objects with a vision-language model, and associates the multi-view detections into 3D instances through Gaussian set matching, which measures the overlap between the Gaussian sets that render each detection. Each source then receives generated audio that a standard object-based audio engine spatializes in real time at arbitrary listener poses. We further propose two spatial-consistency metrics, one testing whether rendered audio responds consistently to listener motion and one testing whether the claimed sources are supported by views held out from their placement. On a curated set of generated 3DGS worlds and on 3DGS scenes generated from real-world 360-degree captures, Scene2Sound preserves the audio quality of strong per-viewpoint baselines while remaining spatially consistent where per-viewpoint and single-panorama pipelines do not, and a user study confirms the perceptual benefit. Project page: https://masaki-lmd.github.io/scene2sound/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。