arXiv:2504.12796cs.MMcs.SD2025-04综述被引 2

梳理音乐与多模态数据的三类交互方式,助力智能音乐研究

A Survey on Cross-Modal Interaction Between Music and Multimodal Data

  • 按音乐主导、响应或双向分三类跨模态交互
  • 系统整理音乐数据集与评估指标,提供研究基准
  • 适合想拓展音乐计算边界的学者参考

多模态学习推动了多个领域的创新,尤其在音乐领域。通过实现更自然的交互体验和增强沉浸感,它不仅降低了音乐技术的使用门槛,还提升了整体吸引力。本文旨在全面回顾与音乐相关的多模态任务,阐明音乐如何促进多模态学习,并为希望拓展计算音乐边界的研究者提供洞见。不同于文本和图像的语义或视觉直观性,音乐主要通过听觉感知与人互动,其数据表示本身较不直观。因此,本文首先介绍音乐的数据表示,并概述常用音乐数据集。随后,将音乐与多模态数据的跨模态交互分为三类:音乐驱动的跨模态交互、以音乐为导向的跨模态交互,以及双向音乐跨模态交互。针对每一类别,系统梳理相关子任务的发展脉络,分析现有局限并探讨新兴趋势。此外,全面总结音乐相关多模态任务所用数据集与评估指标,为未来研究提供基准参考。最后,讨论当前音乐跨模态交互面临的挑战,并提出潜在研究方向。

原文摘要 · Abstract (English)

Multimodal learning has driven innovation across various industries, particularly in the field of music. By enabling more intuitive interaction experiences and enhancing immersion, it not only lowers the entry barriers to the music but also increases its overall appeal. This survey aims to provide a comprehensive review of multimodal tasks related to music, outlining how music contributes to multimodal learning and offering insights for researchers seeking to expand the boundaries of computational music. Unlike text and images, which are often semantically or visually intuitive, music primarily interacts with humans through auditory perception, making its data representation inherently less intuitive. Therefore, this paper first introduces the representations of music and provides an overview of music datasets. Subsequently, we categorize cross-modal interactions between music and multimodal data into three types: music-driven cross-modal interactions, music-oriented cross-modal interactions, and bidirectional music cross-modal interactions. For each category, we systematically trace the development of relevant sub-tasks, analyze existing limitations, and discuss emerging trends. Furthermore, we provide a comprehensive summary of datasets and evaluation metrics used in multimodal tasks related to music, offering benchmark references for future research. Finally, we discuss the current challenges in cross-modal interactions involving music and propose potential directions for future research.

音乐生成多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。