arXiv:2506.23009cs.CV2025-06被引 8

首个音乐谱视觉理解数据集,推动多模态大模型读懂乐谱。

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

  • 构建合成乐谱数据集,涵盖音高、时值、和弦等结构化标注。
  • 现有多模态模型在乐谱理解上表现差,新模型比GPT系提升显著。
  • 适合研究音乐生成、跨模态理解的学者与开发者使用。

多模态大语言模型在自然图像、图文文档和图形设计方面已取得显著视觉推理能力,但对乐谱的理解仍处于探索阶段。为填补这一空白,我们提出MusiXQA,首个全面评估和推进多模态大模型音乐谱理解的基准数据集。该数据集采用MusiXTeX生成高质量合成乐谱,包含音高、时值、和弦、谱号、调号/拍号及文本的结构化标注,支持多样化的视觉问答任务。通过大规模评估,我们揭示了当前先进多模态模型在该领域的显著局限性。除基准测试外,我们还开发了在该数据集上微调的Phi-3-MusiX模型,性能显著优于基于GPT的方法。所提出的数据集与模型为未来多模态大模型在音乐谱理解方向的发展奠定了基础。代码、数据与模型将在论文接受后公开。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this gap, we introduce MusiXQA, the first comprehensive dataset for evaluating and advancing MLLMs in music sheet understanding. MusiXQA features high-quality synthetic music sheets generated via MusiXTeX, with structured annotations covering note pitch and duration, chords, clefs, key/time signatures, and text, enabling diverse visual QA tasks. Through extensive evaluations, we reveal significant limitations of current state-of-the-art MLLMs in this domain. Beyond benchmarking, we developed Phi-3-MusiX, an MLLM fine-tuned on our dataset, achieving significant performance gains over GPT-based methods. The proposed dataset and model establish a foundation for future advances in MLLMs for music sheet understanding. Code, data, and model will be released upon acceptance.

音乐理解多模态大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。