首提360度空间可控音视频生成,精准控制视角内容。
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
- 用全景显著图、符号距离图和场景描述三重控制信号
- 生成与环境上下文一致的视角化音视频内容
- 适合虚拟现实、沉浸式体验等需要精准视角控制的场景
扩散模型推动了声音视频生成的发展,但现有方法难以从大范围沉浸式360°环境中精细控制生成特定视角的内容,限制了对镜头外事件感知的音视频体验。据我们所知,这是首个提出可控音视频生成框架的工作,填补该空白。具体而言,我们引入一种扩散模型,结合三类来自完整360°空间的强条件信号:全景显著图用于识别兴趣区域,边界框感知的符号距离图用于定义目标视角,以及对整个场景的描述性标题。通过整合这些控制信号,模型生成具有空间感知能力的视角视频与音频,其内容受未直接可见环境上下文的协同影响,实现了真实沉浸式音视频生成所需的关键可控性。我们展示了音视频示例,验证了该框架的有效性。
原文摘要 · Abstract (English)
The generation of sounding videos has seen significant advancements with the advent of diffusion models. However, existing methods often lack the fine-grained control needed to generate viewpoint-specific content from larger, immersive 360-degree environments. This limitation restricts the creation of audio-visual experiences that are aware of off-camera events. To the best of our knowledge, this is the first work to introduce a framework for controllable audio-visual generation, addressing this unexplored gap. Specifically, we propose a diffusion model by introducing a set of powerful conditioning signals derived from the full 360-degree space: a panoramic saliency map to identify regions of interest, a bounding-box-aware signed distance map to define the target viewpoint, and a descriptive caption of the entire scene. By integrating these controls, our model generates spatially-aware viewpoint videos and audios that are coherently influenced by the broader, unseen environmental context, introducing a strong controllability that is essential for realistic and immersive audio-visual generation. We show audiovisual examples proving the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。