arXiv:2409.06135cs.SDcs.CV2024-09被引 10

用画图和音量信号控制视频生成配音,让声音更贴合画面。

Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis

  • 通过手绘掩码和音量信号实现多指令控制
  • 在两个大规模数据集上达到当前最佳性能
  • 适合需要精准声画同步的影视音效生成

Foley 指电影制作中为无声影像添加日常音效以增强听觉体验。视频到音频(V2A)作为自动 Foley 的一种,面临音画同步的固有挑战,包括输入视频与生成音频的内容一致性,以及时间与音量属性的对齐问题。为此,我们构建了一个可控的视频到音频合成模型——Draw an Audio,支持通过绘制掩码和音量信号输入多指令。为确保生成音频与目标视频内容一致,引入掩码注意力模块(MAM),利用掩码视频指令使模型聚焦于感兴趣区域。同时,设计时间-音量模块(TLM),使用辅助音量信号确保生成声音在音量和时间维度上与视频对齐。此外,我们扩展了大规模 V2A 数据集 VGGSound-Caption,增加了标注的描述性提示。在两个大规模 V2A 数据集上的大量实验验证,Draw an Audio 达到了当前最优水平。

原文摘要 · Abstract (English)

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents inherent challenges related to audio-visual synchronization. These challenges encompass maintaining the content consistency between the input video and the generated audio, as well as the alignment of temporal and loudness properties within the video. To address these issues, we construct a controllable video-to-audio synthesis model, termed Draw an Audio, which supports multiple input instructions through drawn masks and loudness signals. To ensure content consistency between the synthesized audio and target video, we introduce the Mask-Attention Module (MAM), which employs masked video instruction to enable the model to focus on regions of interest. Additionally, we implement the Time-Loudness Module (TLM), which uses an auxiliary loudness signal to ensure the synthesis of sound that aligns with the video in both loudness and temporal dimensions. Furthermore, we have extended a large-scale V2A dataset, named VGGSound-Caption, by annotating caption prompts. Extensive experiments on challenging benchmarks across two large-scale V2A datasets verify Draw an Audio achieves the state-of-the-art. Project page: https://yannqi.github.io/Draw-an-Audio/.

视频生成音画同步可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。