让AI读懂乐谱与演奏音频,实现音乐的深度交互理解
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
- 用乐谱图像和演奏音频生成结构化符号表示,增强语言模型的音乐感知能力
- 在多模态音乐理解任务上显著超越现有模型,尤其在乐理推理和演奏分析中表现突出
- 适合音乐信息检索、智能作曲辅助等需要精准音乐理解的场景
尽管多模态大语言模型(MLLMs)取得进展,其对音乐的理解与交互能力仍受限。音乐理解需基于乐谱符号与表达性演奏音频进行具身推理,而通用型MLLM因感知基础不足难以胜任。我们提出MuseAgent,一个以音乐为中心的多模态智能体,通过光学乐谱识别与自动乐谱转录模块,将乐谱图像与演奏音频转化为结构化符号表示,使语言模型能对精细音乐内容进行多步推理与交互。为进一步系统评估音乐理解能力,我们构建了MuseBench基准,涵盖音乐理论推理、乐谱解读与演奏级分析,跨文本、图像、音频模态。实验表明,现有MLLM在这些任务上表现不佳,而MuseAgent实现显著提升,凸显结构化多模态具身推理在交互式音乐理解中的关键作用。
原文摘要 · Abstract (English)
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio, which general-purpose MLLMs often fail to handle due to insufficient perceptual grounding. We introduce MuseAgent, a music-centric multimodal agent that augments language models with structured symbolic representations derived from sheet music images and performance audio. By integrating optical music recognition and automatic music transcription modules, MuseAgent enables multi-step reasoning and interaction over fine-grained musical content. To systematically evaluate music understanding capabilities, we further propose MuseBench, a benchmark covering music theory reasoning, score interpretation, and performance-level analysis across text, image, and audio modalities. Experiments show that existing MLLMs perform poorly on these tasks, while MuseAgent achieves substantial improvements, highlighting the importance of structured multimodal grounding for interactive music understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。