arXiv:2506.15759cs.SDcs.MM2025-06AAAI被引 4

让4D动态场景自带真实空间音效,无需训练即可实现沉浸式视听体验。

Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration

  • 通过视觉定位追踪声源,将单声道音频转为随视角变化的空间音频
  • 基于物理仿真生成随时间与视角变化的逼真空间音效
  • 无需训练,直接适配已生成的4D场景,适合虚拟现实应用

4D生成技术已能合成动态3D场景的逼真画面,但几乎所有方法均忽略与之匹配的空间音频,限制了沉浸感。为此,我们提出Sonic4D框架,实现4D场景的沉浸式空间音频生成。首先,利用预训练模型从单目视频中重建4D场景及对应单声道音频;其次,通过像素级视觉对齐策略,在4D场景中定位并跟踪声源,估计其在不同时刻的3D坐标;最后,基于声源位置,使用物理基仿真生成随视角和时间变化的合理空间音频。大量实验表明,该方法以无训练方式生成与合成4D场景一致的真实空间音频,显著提升用户沉浸体验。相关音视频示例见https://x-drunker.github.io/Sonic4D-project-page。

原文摘要 · Abstract (English)

Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook the generation of spatial audio aligned with the corresponding 4D scenes, posing a significant limitation to truly immersive audiovisual experiences. To mitigate this issue, we propose Sonic4D, a novel framework that enables spatial audio generation for immersive exploration of 4D scenes. Specifically, our method is composed of three stages: 1) To capture both the dynamic visual content and raw auditory information from a monocular video, we first employ pre-trained expert models to generate the 4D scene and its corresponding monaural audio. 2) Subsequently, to transform the monaural audio into spatial audio, we localize and track the sound sources within the 4D scene, where their 3D spatial coordinates at different timestamps are estimated via a pixel-level visual grounding strategy. 3) Based on the estimated sound source locations, we further synthesize plausible spatial audio that varies across different viewpoints and timestamps using physics-based simulation. Extensive experiments have demonstrated that our proposed method generates realistic spatial audio consistent with the synthesized 4D scene in a training-free manner, significantly enhancing the immersive experience for users. Generated audio and video examples are available at https://x-drunker.github.io/Sonic4D-project-page.

4D生成空间音频沉浸体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。