用视频生成同步音频,兼顾语义与节奏一致性。
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
- 通过掩码建模对齐音视频编码器,恢复被遮蔽的音频。
- 动态条件流根据视频时序特征生成对应音频段。
- 在标准数据集上性能超越现有方法,适合音视频生成任务。
基于视频输入的协同音频生成通常需要严格的音视频(AV)对齐,即生成的音频段在语义和节奏上均需与视频帧匹配。以往研究采用两阶段设计:先通过对比学习对齐音视频编码器,再由视频表征引导音频生成。我们发现,对比学习和全局视频引导虽能对齐整体音视频语义,但限制了时间上的节奏同步。为此,本文提出FoleyFlow,首先通过掩码建模训练对单模态音视频编码器进行对齐,其中被遮蔽的音频段在对应视频段的指导下被恢复。训练完成后,仅使用单模态数据预训练的音视频编码器实现了语义与节奏的一致性对齐。随后,我们构建动态条件流用于最终音频生成。基于高效的速度流生成框架,该动态条件流利用随时间变化的视频特征作为动态条件,指导对应音频段的生成。为此,我们在掩码音视频对齐过程中提取出一致的语义与节奏表征,并将其用于时序化音频生成。在标准基准上的评估显示,我们的音频结果在多个指标上显著优于现有方法。优越性能表明FoleyFlow能有效生成与各类视频序列在语义和节奏上均高度一致的协同音频。
原文摘要 · Abstract (English)
Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies leverage a two-stage design where the AV encoders are firstly aligned via contrastive learning, then the encoded video representations guide the audio generation process. We observe that both contrastive learning and global video guidance are effective in aligning overall AV semantics while limiting temporally rhythmic synchronization. In this work, we propose FoleyFlow to first align unimodal AV encoders via masked modeling training, where the masked audio segments are recovered under the guidance of the corresponding video segments. After training, the AV encoders which are separately pretrained using only unimodal data are aligned with semantic and rhythmic consistency. Then, we develop a dynamic conditional flow for the final audio generation. Built upon the efficient velocity flow generation framework, our dynamic conditional flow utilizes temporally varying video features as the dynamic condition to guide corresponding audio segment generations. To this end, we extract coherent semantic and rhythmic representations during masked AV alignment, and use this representation of video segments to guide audio generation temporally. Our audio results are evaluated on the standard benchmarks and largely surpass existing results under several metrics. The superior performance indicates that FoleyFlow is effective in generating coordinated audios that are both semantically and rhythmically coherent to various video sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。