arXiv:2503.21254cs.CVcs.AI2025-03综述被引 8

系统梳理视觉生成音乐的研究进展与挑战

Vision-to-Music Generation: A Survey

  • 按输入类型和输出形式分类,分析技术特点与核心难点
  • 总结现有方法架构,涵盖视频、动作、图像到音乐的生成路径
  • 适合关注多模态生成与创意产业应用的研究者

视觉生成音乐(包括视频到音乐、图像到音乐)是多模态人工智能的重要分支,在影视配乐、短视频创作和舞蹈音乐合成等领域具有广阔应用前景。然而,由于其内部结构复杂且难以建模视频中动态关系,该领域研究仍处于初步阶段,相较于文本与图像等模态发展滞后。现有综述多聚焦通用音乐生成,缺乏对视觉到音乐生成的全面讨论。本文系统回顾了视觉生成音乐的研究进展,首先分析三类输入(通用视频、人体动作视频、图像)与两类输出(符号化音乐、音频音乐)的技术特征与核心挑战;其次从架构角度总结现有生成方法;详尽梳理常用数据集与评估指标;最后探讨当前瓶颈与未来发展方向。我们希望本综述能推动该领域学术与工业创新,并持续维护开源项目:https://github.com/wzk1015/Awesome-Vision-to-Music-Generation。

原文摘要 · Abstract (English)

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and dance music synthesis. However, compared to the rapid development of modalities like text and images, research in vision-to-music is still in its preliminary stage due to its complex internal structure and the difficulty of modeling dynamic relationships with video. Existing surveys focus on general music generation without comprehensive discussion on vision-to-music. In this paper, we systematically review the research progress in the field of vision-to-music generation. We first analyze the technical characteristics and core challenges for three input types: general videos, human movement videos, and images, as well as two output types of symbolic music and audio music. We then summarize the existing methodologies on vision-to-music generation from the architecture perspective. A detailed review of common datasets and evaluation metrics is provided. Finally, we discuss current challenges and promising directions for future research. We hope our survey can inspire further innovation in vision-to-music generation and the broader field of multimodal generation in academic research and industrial applications. To follow latest works and foster further innovation in this field, we are continuously maintaining a GitHub repository at https://github.com/wzk1015/Awesome-Vision-to-Music-Generation.

多模态生成视觉音乐综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。