系统梳理视频生成配乐的主流方法与挑战
A Comprehensive Survey on Generative AI for Video-to-Music Generation
- 按输入构建、条件机制、音乐生成框架分类现有方法
- 细分为视频与音乐模态,揭示不同类别对设计的影响
- 总结数据集与评估指标,指出领域现存难题
视频生成配乐的快速发展得益于多模态生成模型的兴起。然而,该领域尚缺乏全面的文献综述。本文系统回顾了基于深度生成AI的视频到音乐生成技术,聚焦三个核心组件:条件输入构建、条件机制设计和音乐生成框架。我们根据各组件的设计对现有方法进行分类,阐明不同策略的作用。在此之前,我们对视频与音乐模态进行了细粒度划分,展示不同类别如何影响生成流程中的组件设计。此外,本文总结了可用的多模态数据集与评估指标,并指出现有挑战。
原文摘要 · Abstract (English)
The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this paper presents a comprehensive review of video-to-music generation using deep generative AI techniques, focusing on three key components: conditioning input construction, conditioning mechanism, and music generation frameworks. We categorize existing approaches based on their designs for each component, clarifying the roles of different strategies. Preceding this, we provide a fine-grained categorization of video and music modalities, illustrating how different categories influence the design of components within the generation pipelines. Furthermore, we summarize available multimodal datasets and evaluation metrics while highlighting ongoing challenges in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。