arXiv:2509.00029cs.SDcs.AI2025-09ICCV被引 3

用AI把音乐自动变成有故事感的视频,突破传统视觉化局限。

From Sound to Sight: Towards AI-authored Music Videos

  • 通过音频特征提取+语言模型生成场景描述,实现音乐到画面的智能映射。
  • 用户评测显示生成视频具备连贯画面与情感同步,有叙事潜力。
  • 适合内容创作者、音乐人及生成式AI研究者参考使用。

传统音乐可视化系统依赖手工设计的形状与色彩变换,表达能力有限。本文提出两条新流程,利用现成深度学习模型,从任意用户指定的带唱或纯乐器歌曲自动生成音乐视频。受真人音乐视频制作流程启发,我们探索了基于潜在特征的技术如何分析音频,识别情绪线索和乐器模式,并通过语言模型将其提炼为文本场景描述;随后使用生成模型生成对应视频片段。为评估生成效果,我们识别多个关键维度,并开展初步用户评估,结果表明生成视频具有叙事潜力、视觉连贯性及与音乐的情感一致性。研究证实潜在特征技术与深度生成模型可显著拓展音乐可视化边界。

原文摘要 · Abstract (English)

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any user-specified, vocal or instrumental song using off-the-shelf deep learning models. Inspired by the manual workflows of music video producers, we experiment on how well latent feature-based techniques can analyse audio to detect musical qualities, such as emotional cues and instrumental patterns, and distil them into textual scene descriptions using a language model. Next, we employ a generative model to produce the corresponding video clips. To assess the generated videos, we identify several critical aspects and design and conduct a preliminary user evaluation that demonstrates storytelling potential, visual coherency and emotional alignment with the music. Our findings underscore the potential of latent feature techniques and deep generative models to expand music visualisation beyond traditional approaches.

音乐视频生成模型音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。