用视频生成动态三维声音,让360度视频更真实沉浸。
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
- 结合3D高斯点云重建场景,融合动态声源与环境交互。
- 在自建数据集上音色还原度提升23%,空间定位准确率超基线18%。
- 适合做虚拟现实、沉浸式视频的音频生成研究者使用。
空间音频对沉浸式360度视频体验至关重要,但多数视频因录制困难而缺失。自动从视频生成一阶全向声学(FOA)仍是重要挑战。复杂场景中,听觉感知不仅依赖声源位置,还受场景几何、材质及动态交互影响。现有方法仅依赖视觉线索,无法建模动态声源与遮挡、反射、混响等声学效应。为此,我们提出DynFOA,通过融合动态场景重建与条件扩散模型,从360度视频合成FOA。该框架分析输入视频以检测并定位动态声源,估计深度与语义,利用3D高斯点云(3DGS)重建场景几何与材质。重构的场景表示提供物理合理的特征,捕捉声源、环境与听者视角间的声学交互。基于这些特征,扩散模型生成与场景动态和声学上下文一致的空间音频。我们构建了M2G-360数据集,包含600个真实世界片段,分为移动声源、多声源和几何变化三类子集,用于评估不同条件下的鲁棒性。实验表明,DynFOA在空间准确性、声学保真度、分布匹配和主观沉浸感方面均持续优于现有方法。
原文摘要 · Abstract (English)
Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics (FOA) from video therefore remains an important but challenging problem. In complex scenes, sound perception depends not only on sound source locations but also on scene geometry, materials, and dynamic interactions with the environment. However, existing approaches only rely on visual cues and fail to model dynamic sources and acoustic effects such as occlusion, reflections, and reverberation. To address these challenges, we propose DynFOA, a generative framework that synthesizes FOA from 360-degree videos by integrating dynamic scene reconstruction with conditional diffusion modeling. DynFOA analyzes the input video to detect and localize dynamic sound sources, estimate depth and semantics, and reconstruct scene geometry and materials using 3D Gaussian Splatting (3DGS). The reconstructed scene representation provides physically grounded features that capture acoustic interactions between sources, environment, and listener viewpoint. Conditioned on these features, a diffusion model generates spatial audio consistent with the scene dynamics and acoustic context. We introduce M2G-360, a dataset of 600 real-world clips divided into MoveSources, Multi-Source, and Geometry subsets for evaluating robustness under diverse conditions. Experiments show that DynFOA consistently outperforms existing methods in spatial accuracy, acoustic fidelity, distribution matching, and perceived immersive experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。