arXiv:2509.23350cs.SDcs.AI2025-09被引 9

首个评测大模型符号音乐理解与指令遵循能力的开源基准

ABC-Eval: Benchmarking Large Language Models on Symbolic Music Understanding and Instruction Following

  • 构建涵盖10类任务的符号音乐评测集,覆盖从语法到复杂推理
  • 7个主流大模型在任务中表现普遍不佳,暴露出处理音乐符号的短板
  • 适合音乐智能、人机交互与多模态研究者参考

随着大语言模型的发展,基于文本的符号音乐任务日益重要。尽管符号音乐广泛应用于生成任务,但大模型在理解和推理符号音乐方面仍缺乏系统探索。为此,我们提出ABC-Eval,首个专注于文本式ABC记谱法理解与指令遵循能力的开源基准。该基准包含1,086个测试样本,覆盖10个子任务,从基础音乐语法理解到复杂序列级推理均有涉及。多样化的任务设计对模型处理符号音乐的能力构成严峻挑战。我们在七个先进大模型上进行了评估,结果揭示了现有模型在符号音乐处理上的显著局限性。此外,各基线模型在不同子任务间表现稳定,验证了本基准的可靠性。

原文摘要 · Abstract (English)

As large language models continue to develop, the feasibility and significance of text-based symbolic music tasks have become increasingly prominent. While symbolic music has been widely used in generation tasks, LLM capabilities in understanding and reasoning about symbolic music remain largely underexplored. To address this gap, we propose ABC-Eval, the first open-source benchmark dedicated to the understanding and instruction-following capabilities in text-based ABC notation scores. It comprises 1,086 test samples spanning 10 sub-tasks, covering scenarios from basic musical syntax comprehension to complex sequence-level reasoning. Such a diverse scope poses substantial challenges to models' ability to handle symbolic music tasks. We evaluated seven state-of-the-art LLMs on ABC-Eval, and the results reveal notable limitations in existing models' symbolic music processing capabilities. Furthermore, the consistent performance of individual baselines across different sub-tasks supports the reliability of our benchmark.

符号音乐大模型评测指令遵循ABC记谱法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。