arXiv:2409.07450cs.MMcs.CV2024-09被引 24

用网页视频自动生成匹配的背景音乐,效果更真实多样。

VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

  • 通过语义对齐机制让音乐与视频内容高阶特征匹配
  • 在220万样本数据集上训练,生成音乐更自然且多样性更高
  • 适合音乐生成、跨模态创作研究者使用

我们提出一种从视频输入生成背景音乐的框架。不同于依赖稀缺符号音乐标注的现有方法,本方法利用大量带有背景音乐的网络视频进行训练,使模型能够学习生成真实且多样的音乐。为此,我们设计了一个生成式视频-音乐Transformer,并引入新颖的语义视频-音乐对齐方案。模型采用联合自回归与对比学习目标,促使生成音乐与视频高层内容保持一致。此外,我们提出一种新的视频节拍对齐机制,使生成音乐节拍与视频低层运动相匹配。为捕捉生成真实背景音乐所需的细粒度视觉线索,我们设计了一种新型时序视频编码器结构,可高效处理密集采样的视频帧。我们在新构建的DISCO-MV数据集上训练,该数据集包含220万组视频-音乐样本,规模远超以往相关数据集。实验表明,本方法在DISCO-MV和MusicCaps数据集上的多项音乐生成评估指标(包括人工评价)均优于现有方法。

原文摘要 · Abstract (English)

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos accompanied by background music. This enables our model to learn to generate realistic and diverse music. To accomplish this goal, we develop a generative video-music Transformer with a novel semantic video-music alignment scheme. Our model uses a joint autoregressive and contrastive learning objective, which encourages the generation of music aligned with high-level video content. We also introduce a novel video-beat alignment scheme to match the generated music beats with the low-level motions in the video. Lastly, to capture fine-grained visual cues in a video needed for realistic background music generation, we introduce a new temporal video encoder architecture, allowing us to efficiently process videos consisting of many densely sampled frames. We train our framework on our newly curated DISCO-MV dataset, consisting of 2.2M video-music samples, which is orders of magnitude larger than any prior datasets used for video music generation. Our method outperforms existing approaches on the DISCO-MV and MusicCaps datasets according to various music generation evaluation metrics, including human evaluation. Results are available at https://genjib.github.io/project_page/VMAs/index.html

视频配乐生成模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。