用视频视觉特征生成匹配语义和节奏的音乐
VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
- 分层视觉特征分别控制音乐语义与节奏
- 在多个指标上超越现有方法,对AI视频也有效
- 适合视频配乐、内容创作等场景
视频到音乐生成在视频制作中具有巨大潜力,要求生成的音乐在语义和节奏上均与视频对齐。这需要强大的音乐生成能力、精细的视频理解以及模态间对应关系的学习机制。本文提出VidMusician,一种基于文本到音乐模型的参数高效视频到音乐生成框架。该方法利用层次化视觉特征实现视频与音乐的语义和节奏对齐:全局视觉特征作为语义条件,局部视觉特征作为节奏线索,分别通过交叉注意力和内部注意力机制融入生成主干。采用两阶段训练,逐步引入语义与节奏特征,结合零初始化与恒等初始化以保留主干原有的音乐生成能力。此外,构建了涵盖宣传片、广告、混剪等多种场景的多样化视频-音乐数据集DVMSet。实验表明,VidMusician在多个评估指标上优于当前最优方法,并在AI生成视频上表现稳健。样例可访问:https://youtu.be/EPOSXwtl1jw。
原文摘要 · Abstract (English)
Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achieving this alignment demands advanced music generation capabilities, sophisticated video understanding, and an efficient mechanism to learn the correspondence between the two modalities. In this paper, we propose VidMusician, a parameter-efficient video-to-music generation framework built upon text-to-music models. VidMusician leverages hierarchical visual features to ensure semantic and rhythmic alignment between video and music. Specifically, our approach utilizes global visual features as semantic conditions and local visual features as rhythmic cues. These features are integrated into the generative backbone via cross-attention and in-attention mechanisms, respectively. Through a two-stage training process, we incrementally incorporate semantic and rhythmic features, utilizing zero initialization and identity initialization to maintain the inherent music-generative capabilities of the backbone. Additionally, we construct a diverse video-music dataset, DVMSet, encompassing various scenarios, such as promo videos, commercials, and compilations. Experiments demonstrate that VidMusician outperforms state-of-the-art methods across multiple evaluation metrics and exhibits robust performance on AI-generated videos. Samples are available at \url{https://youtu.be/EPOSXwtl1jw}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。