用分层注意力实现视频到音乐的精准生成,支持多风格零样本输出。
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
- 通过分层注意力对齐视频与音乐的时空特征,减少冗余信息。
- 在多风格生成和零样本场景下表现优于现有模型。
- 构建大规模视频-音乐数据集,提出新评估指标衡量匹配度。
为视频配乐是重要但具挑战性的任务,近年来自动化音乐生成备受关注。现有方法常因特征对齐不足和数据集有限,难以保证音乐与视频的强关联性及生成多样性。本文提出通用视频到音乐生成模型GVMGen,利用分层注意力在时空维度上提取并对齐视频与音乐特征,保留关键信息同时降低冗余。该方法具备高度灵活性,可从不同视频输入生成多种风格音乐,甚至在零样本场景下仍有效。我们还设计了评估模型及两项新客观指标以衡量视频-音乐匹配度,并构建了一个涵盖多种类型视频-音乐对的大规模数据集。实验表明,GVMGen在音乐-视频对应性、生成多样性及应用普适性方面均超越先前模型。
原文摘要 · Abstract (English)
Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods and insufficient datasets. In this study, we present General Video-to-Music Generation model (GVMGen), designed for generating high-related music to the video input. Our model employs hierarchical attentions to extract and align video features with music in both spatial and temporal dimensions, ensuring the preservation of pertinent features while minimizing redundancy. Remarkably, our method is versatile, capable of generating multi-style music from different video inputs, even in zero-shot scenarios. We also propose an evaluation model along with two novel objective metrics for assessing video-music alignment. Additionally, we have compiled a large-scale dataset comprising diverse types of video-music pairs. Experimental results demonstrate that GVMGen surpasses previous models in terms of music-video correspondence, generative diversity, and application universality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。