综述多模态音乐生成技术,涵盖文本、图像、视频等引导作曲的前沿进展。
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
- 按模态分类梳理音乐生成系统,分析跨模态表示与对齐方法
- 指出当前数据集规模小、评估体系不统一等核心挑战
- 适合关注创意生成、跨模态融合的研究者和应用开发者
多模态音乐生成利用文本、图像、视频及乐谱、音频等多种模态作为引导,是新兴研究方向且应用广泛。本文从模态视角综述该领域,涵盖模态表征、多模态数据对齐及其在音乐生成中的应用。同时讨论现有数据集与评估方法。主要挑战包括有效的多模态融合、大规模综合数据集的缺乏以及系统性评估体系的缺失。最后展望未来研究方向,聚焦创造力提升、生成效率优化、多模态对齐机制与评价标准建设。
原文摘要 · Abstract (English)
Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music generation systems from the perspective of modalities. The review covers modality representation, multi-modal data alignment, and their utilization to guide music generation. Current datasets and evaluation methods are also discussed. Key challenges in this area include effective multi-modal integration, large-scale comprehensive datasets, and systematic evaluation methods. Finally, an outlook on future research directions is provided, focusing on creativity, efficiency, multi-modal alignment, and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。