让视频生成有空间感的立体声,还能对应具体物体。
StereoFoley: Object-Aware Stereo Audio Generation from Video
- 用视频分析+动态音量和声像控制生成立体声
- 在合成数据上微调后实现物体与声音精准对应
- 首个端到端的立体物体感知音视频生成框架
我们提出 StereoFoley,一个从视频生成语义对齐、时间同步且空间准确的48 kHz立体声音频的框架。尽管现有视频转音频模型在语义和时间一致性上表现良好,但大多仅支持单声道或无法实现物体感知的立体声,受限于缺乏专业混音、空间准确的视频-音频数据集。首先,我们构建基础模型,在语义准确性和同步性上达到当前最优水平。其次,为突破数据限制,提出合成数据生成流程,结合视频分析、目标追踪与音频合成,通过动态声像偏移和距离相关音量控制,实现空间精确的物体感知声音。最后,基于该合成数据微调模型,获得清晰的物体-音频对应关系。由于缺乏标准评估指标,我们引入立体声物体感知度量,并结合人工听觉测试;两者结果趋势一致。本工作建立了首个端到端的立体物体感知视频转音频生成框架,填补了该领域的关键空白。
原文摘要 · Abstract (English)
We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic and temporal fidelity, they largely remain limited to mono or fail to deliver object-aware stereo imaging, constrained by the lack of professionally mixed, spatially accurate video-to-audio datasets. First, we develop a base model that generates stereo audio from video, achieving performance on par with state-of-the-art V2A models in both semantic accuracy and synchronization. Next, to overcome dataset limitations, we introduce a synthetic data generation pipeline that combines video analysis, object tracking, and audio synthesis with dynamic panning and distance-based loudness controls, enabling spatially accurate object-aware sound. Finally, we fine-tune the base model on this synthetic dataset, yielding clear object-audio correspondence. Since no established metrics exist, we introduce a stereo object-awareness metric and report it alongside a human listening study; the two evaluations exhibit consistent trends. This work establishes the first end-to-end framework for stereo object-aware video-to-audio generation, addressing a critical gap in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。