从视频生成高保真长音频,8步即完成,同步精准。
SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
- 用掩码扩散机制实现视频到音频的条件生成
- 仅需8步采样即可生成高质量音频,接近实时
- 适合音效设计、影视配音等专业音频合成场景
我们提出SALSA-V,一种多模态视频转音频生成模型,可从无声视频内容合成高度同步、高保真的长时音频。该方法引入掩码扩散目标,支持音频条件生成,并实现无长度限制音频序列的无缝合成。通过在训练中加入捷径损失,模型可在仅8次采样步骤内快速生成高质量音频样本,为无需微调或重训练的近实时应用铺平道路。定量评估与人工听觉实验均表明,SALSA-V在音视频对齐与同步性上显著优于现有最先进方法。此外,训练中采用随机掩码策略,使模型能匹配参考音频的频谱特性,拓展其在专业音效制作(如拟音、声音设计)中的适用性。
原文摘要 · Abstract (English)
We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of unconstrained length audio sequences. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantitative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。