用视频生成动态360度空间音频,还原声音与环境互动效果。
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
- 结合3D高斯泼溅重建场景,捕捉声源与环境交互特征。
- 在M2G-360数据集上,空间精度和沉浸感显著优于现有方法。
- 适合需要真实感空间音频的虚拟现实与全景视频应用。
空间音频对沉浸式360度视频体验至关重要,但多数视频因录制困难缺乏空间音频。自动从视频生成第一阶球谐音频(FOA)仍是重要挑战。复杂场景中,听觉感知不仅依赖声源位置,还受场景几何、材质及动态交互影响。现有方法仅依赖视觉线索,无法建模动态声源与遮挡、反射、混响等声学效应。为此,我们提出DynFOA,通过融合动态场景重建与条件扩散模型,从360度视频生成FOA。该方法分析输入视频,检测并定位动态声源,估计深度与语义,利用3D高斯泼溅(3DGS)重建场景几何与材质。重建后的场景表示提供物理基础特征,捕捉声源、环境与听者视角间的声学交互。以这些特征为条件,扩散模型生成与场景动态和声学上下文一致的空间音频。我们构建了M2G-360数据集,包含600个真实世界片段,分为移动声源、多声源和几何变化三类,用于评估复杂条件下的鲁棒性。实验表明,DynFOA在空间准确性、声学保真度、分布匹配性和感知沉浸感方面持续优于现有方法。
原文摘要 · Abstract (English)
Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics (FOA) from video therefore remains an important but challenging problem. In complex scenes, sound perception depends not only on sound source locations but also on scene geometry, materials, and dynamic interactions with the environment. However, existing approaches only rely on visual cues and fail to model dynamic sources and acoustic effects such as occlusion, reflections, and reverberation. To address these challenges, we propose DynFOA, a generative framework that synthesizes FOA from 360-degree videos by integrating dynamic scene reconstruction with conditional diffusion modeling. DynFOA analyzes the input video to detect and localize dynamic sound sources, estimate depth and semantics, and reconstruct scene geometry and materials using 3D Gaussian Splatting (3DGS). The reconstructed scene representation provides physically grounded features that capture acoustic interactions between sources, environment, and listener viewpoint. Conditioned on these features, a diffusion model generates spatial audio consistent with the scene dynamics and acoustic context. We introduce M2G-360, a dataset of 600 real-world clips divided into MoveSources, Multi-Source, and Geometry subsets for evaluating robustness under diverse conditions. Experiments show that DynFOA consistently outperforms existing methods in spatial accuracy, acoustic fidelity, distribution matching, and perceived immersive experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。