一张图生成可漫游3D音视频场景,让视觉世界有了声音
SonoWorld: From One Image to a 3D Audio-Visual Scene
- 从单张图像生成360°全景并构建可导航3D场景
- 通过语言引导放置声源锚点,实现与场景结构对齐的空间音频
- 支持音视频协同的零样本声学学习与空间声源分离
视觉场景生成已能将单张图像拓展为可探索的3D世界,但缺乏声音使沉浸感不完整。我们提出Image2AVScene任务,即从单张图像生成3D音视频场景,并推出首个解决该问题的框架SonoWorld。该流程首先外推生成360°全景图,将其提升为可导航的3D场景,再基于语言提示放置音源锚点,最后渲染包含点声源、面声源和环境声的ambisonics音频,实现与场景几何与语义一致的空间音频。在新构建的真实世界数据集和受控用户研究中,定量评估验证了方法的有效性。除自由视角音视频渲染外,还展示了其在单次样本声学学习与音视频空间声源分离中的应用。项目主页:https://humathe.github.io/sonoworld/
原文摘要 · Abstract (English)
Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio-visual scene from a single image, and present SonoWorld, the first framework to tackle this challenge. From one image, our pipeline outpaints a 360° panorama, lifts it into a navigable 3D scene, places language-guided sound anchors, and renders ambisonics for point, areal, and ambient sources, yielding spatial audio aligned with scene geometry and semantics. Quantitative evaluations on a newly curated real-world dataset and a controlled user study confirm the effectiveness of our approach. Beyond free-viewpoint audio-visual rendering, we also demonstrate applications to one-shot acoustic learning and audio-visual spatial source separation. Project website: https://humathe.github.io/sonoworld/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。