让音乐模型想象匹配的场景,提升视频配乐生成效果
MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation
- 用音乐与视频对构建3371对数据,训练跨模态场景想象能力
- 相比原模型,新模型生成的场景描述更贴合音乐氛围
- 适合做音乐驱动的视频创作与智能配乐系统研发
人类听音乐时能想象出相应的情境和氛围,如慢调哀伤的音乐可能唤起悲伤画面,欢快旋律则联想到庆祝场景。本文探究音乐语言模型(如MU-LLaMA)是否具备类似能力,即音乐场景想象(MSI),该任务需融合视频与音乐的跨模态信息。为超越仅关注音乐元素的传统音乐描述模型,本文提出MusiScene——一个旨在生成与音乐相匹配场景的音乐描述模型。本工作包含三部分:(1) 构建包含3,371对视频-音频样本的大规模跨模态数据集;(2) 在音乐理解版LLaMA基础上微调,实现MSI任务,得到MusiScene模型;(3) 通过全面评估证明,MusiScene在生成上下文相关描述方面优于原始MU-LLaMA。进一步利用生成的MSI描述增强文本到视频背景音乐生成(VBMG)性能。
原文摘要 · Abstract (English)
Humans can imagine various atmospheres and settings when listening to music, envisioning movie scenes that complement each piece. For example, slow, melancholic music might evoke scenes of heartbreak, while upbeat melodies suggest celebration. This paper explores whether a Music Language Model, e.g. MU-LLaMA, can perform a similar task, called Music Scene Imagination (MSI), which requires cross-modal information from video and music to train. To improve upon existing music captioning models which focusing solely on musical elements, we introduce MusiScene, a music captioning model designed to imagine scenes that complement each music. In this paper, (1) we construct a large-scale video-audio caption dataset with 3,371 pairs, (2) we finetune Music Understanding LLaMA for the MSI task to create MusiScene, and (3) we conduct comprehensive evaluations and prove that our MusiScene is more capable of generating contextually relevant captions compared to MU-LLaMA. We leverage the generated MSI captions to enhance Video Background Music Generation (VBMG) from text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。