arXiv:2503.11190cs.SDcs.AI2025-03中稿 · NAACL被引 2

用音乐生成视频描述,让机器理解音乐并说出画面内容。

Cross-Modal Learning for Music-to-Music-Video Description Generation

  • 基于Music4All构建音乐到视频描述数据集,融合音视频信息
  • 微调多模态模型后,音乐可准确生成有意义的视频描述
  • 发现特定音乐属性对描述质量影响大,适合多媒体生成研究者

由于音乐与视频模态的本质差异,音乐到音乐视频生成是一项挑战性任务。随着强大的文本到视频扩散模型的发展,通过先解决音乐到视频描述任务,再利用这些模型生成视频成为可行路径。本研究聚焦于视频描述生成任务,提出涵盖训练数据构建和多模态模型微调的完整流程。我们在基于Music4All数据集构建的新音乐到视频描述数据集上,微调现有的预训练多模态模型,该数据集整合了音乐与视觉信息。实验结果表明,音乐表征能有效映射到文本域,实现从音乐输入直接生成有意义的视频描述。我们还识别出数据集构建流程中的关键组件,这些组件显著影响视频描述质量,并指出需重点关注特定音乐属性以提升生成效果。

原文摘要 · Abstract (English)

Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for music-video (MV) generation by first addressing the music-to-MV description task and subsequently leveraging these models for video generation. In this study, we focus on the MV description generation task and propose a comprehensive pipeline encompassing training data construction and multimodal model fine-tuning. We fine-tune existing pre-trained multimodal models on our newly constructed music-to-MV description dataset based on the Music4All dataset, which integrates both musical and visual information. Our experimental results demonstrate that music representations can be effectively mapped to textual domains, enabling the generation of meaningful MV description directly from music inputs. We also identify key components in the dataset construction pipeline that critically impact the quality of MV description and highlight specific musical attributes that warrant greater focus for improved MV description generation.

音乐生成多模态视频描述扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。