arXiv:2409.08628cs.SDcs.MM2024-09被引 7

让视频生成的音频与画面精准同步,尤其适合快节奏动作场景。

Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

  • 用语义对齐和时间同步双适配器,提升音画一致性。
  • 在快节奏动作中显著改善节拍点对齐精度和音频语义完整性。
  • 支持用户调节控制器,灵活控制音频效果,适合影视音效生成。

本研究提出一种新型视频到音频合成框架,解决音频与视频不同步及语义丢失问题。通过引入语义对齐适配器和时间同步适配器,方法显著提升了语义完整性和节拍点对齐精度,尤其在快速动作序列中表现优异。模型基于对比式音视频预训练编码器,利用视频与高质量音频数据进行训练,有效提升生成音频质量。该双适配器设计使用户能通过调节控制器实现对音频语义和节拍效果的精细控制,获得更优结果。大量实验验证了该框架在实现无缝音画对齐方面的有效性。

原文摘要 · Abstract (English)

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic integrity and the precision of beat point synchronization, particularly in fast-paced action sequences. Utilizing a contrastive audio-visual pre-trained encoder, our model is trained with video and high-quality audio data, improving the quality of the generated audio. This dual-adapter approach empowers users with enhanced control over audio semantics and beat effects, allowing the adjustment of the controller to achieve better results. Extensive experiments substantiate the effectiveness of our framework in achieving seamless audio-visual alignment.

音视频同步音频生成视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。