arXiv:2506.20995cs.CVcs.LG2025-06中稿 · ECCV被引 1

分步生成视频配乐,避免重复音效,提升音频真实感。

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

  • 每步生成时用负向引导抑制已有音效,实现渐进式声音创作。
  • 在单一参考数据集上训练,无需昂贵的多参考音视频对。
  • 适合需要精细控制音效合成的影视制作与虚拟内容创作者。

我们提出一种分步视频到音频(V2A)生成方法,实现更精细的生成控制和更真实的音频合成。受传统音效制作流程启发,该方法支持逐步添加由视频引发的多种声音事件。为避免依赖昂贵的多参考音视频数据集,每个生成步骤均采用负向引导的V2A过程,抑制先前已生成轨道中的声音重复。指导模型通过微调预训练的V2A模型,在同一视频非重叠片段的音频对上进行训练,使模型在利用声学上下文的同时保持视觉一致性,并支持使用标准单参考音视频数据集进行训练。客观与主观评估表明,该方法提升了每一步生成声音的可分离性,改善了最终复合音频的整体质量,优于现有基线方法。

原文摘要 · Abstract (English)

We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach enables incremental generation of complementary sounds, allowing users to author multiple sound events induced by a video. To avoid the need for costly multi-reference video-audio datasets, each generation step is formulated as a negatively guided V2A process that discourages duplication of sounds already present in previously generated tracks. The guidance model is trained by finetuning a pre-trained V2A model on audio pairs from non-overlapping segments of the same video, encouraging it to leverage acoustic context while remaining visually grounded, and enabling training with standard single-reference audiovisual datasets. Objective and subjective evaluations demonstrate that our method enhances the separability of generated sounds at each step and improves the overall quality of the final composite audio, outperforming existing baselines. Our project page is available at: https://ahykw.github.io/sbsv2a/.

视频配音分步生成负向引导音频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。