让音频随视频画面动态变化,实现时空同步的立体音效生成
StereoSync: Spatially-Aware Stereo Audio Generation from Video

- 利用深度图与边界框提取空间线索,引导扩散模型生成立体声
- 在游戏场景视频上实现音画时空精准对齐,提升沉浸感
- 基于预训练模型,高效生成高质量音频,适合影视音效创作
尽管音频生成近年来受到广泛关注,但与视频对齐的音频生成仍属未充分探索的领域。为此,我们提出 StereoSync,一种新颖且高效的模型,能够生成在时间上与参考视频同步、在空间上与视觉上下文对齐的音频。StereoSync 通过利用预训练基础模型,在减少大量训练需求的同时保持高质量合成。与以往仅关注时间同步的方法不同,StereoSync 首次将空间感知引入视频对齐音频生成中。具体而言,给定输入视频,该方法从深度图和边界框中提取空间线索,并将其作为交叉注意力条件,嵌入基于扩散的音频生成模型中。这一机制使 StereoSync 能够超越简单同步,生成动态适应视频场景空间结构与运动的立体音频。我们在 Walking The Maps 数据集上进行评估,该数据集由来自视频游戏的动画角色行走于多样化环境中的视频组成。实验结果表明,StereoSync 可实现时间和空间双重对齐,显著推进视频到音频生成的前沿水平,带来更沉浸、更真实的听觉体验。
原文摘要 · Abstract (English)
Although audio generation has been widely studied over recent years, video-aligned audio generation still remains a relatively unexplored frontier. To address this gap, we introduce StereoSync, a novel and efficient model designed to generate audio that is both temporally synchronized with a reference video and spatially aligned with its visual context. Moreover, StereoSync also achieves efficiency by leveraging pretrained foundation models, reducing the need for extensive training while maintaining high-quality synthesis. Unlike existing methods that primarily focus on temporal synchronization, StereoSync introduces a significant advancement by incorporating spatial awareness into video-aligned audio generation. Indeed, given an input video, our approach extracts spatial cues from depth maps and bounding boxes, using them as cross-attention conditioning in a diffusion-based audio generation model. Such an approach allows StereoSync to go beyond simple synchronization, producing stereo audio that dynamically adapts to the spatial structure and movement of a video scene. We evaluate StereoSync on Walking The Maps, a curated dataset comprising videos from video games that feature animated characters walking through diverse environments. Experimental results demonstrate the ability of StereoSync to achieve both temporal and spatial alignment, advancing the state of the art in video-to-audio generation and resulting in a significantly more immersive and realistic audio experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。