arXiv:2601.11556cs.LGcs.AI2026-01被引 2

构建音乐符号推理的组合式检索基准,提升AI理解复杂乐理查询能力。

CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning

论文配图:CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
图 1 · 摘自论文原文
  • 设计可链式分析的多步音乐推理框架,结合符号化操作与推理控制器
  • 在126个真实场景问题上实现5-7%准确率提升,尤其在复杂分析类任务中表现显著
  • 适合音乐信息检索、AI作曲及符号计算方向的研究者使用

自然语言对乐谱的查询常需多步组合式音乐信息检索(MIR),从结构化记谱中提取多个证据并整合以回答问题。现有基准难以全面捕捉此类需求,多聚焦于孤立知识或简化场景。本文提出CSyMR-Bench,一个基于社区讨论和专业考试的真实用户场景下的组合式音乐推理基准,包含126道选择题,每题需对乐谱进行多步原子分析以推导隐含音乐证据。为支持诊断,提供六类查询意图与六类分析维度标签。进一步提出工具增强型检索推理框架CSyMR-Agent,融合ReAct风格控制器与基于music21的确定性符号分析算子。实验表明,工具驱动的组合检索显著优于纯大模型方法,在多个类别中取得5-7%的绝对准确率提升,尤其在分析密集型任务中效果最显著。

原文摘要 · Abstract (English)

Natural language information needs over symbolic music scores rarely reduce to a single step lookup. Many queries require compositional Music Information Retrieval (MIR) that extracts multiple pieces of evidence from structured notation and aggregates them to answer the question. This setting remains challenging for Large Language Models due to the mismatch between natural language intents and symbolic representations, as well as the difficulty of reliably handling long structured contexts. Existing benchmarks only partially capture these retrieval demands, often emphasizing isolated theoretical knowledge or simplified settings. We introduce CSyMR-Bench, a benchmark for compositional MIR in symbolic music reasoning grounded in authentic user scenarios. It contains 126 multiple choice questions curated from community discussions and professional examinations, where each item requires chaining multiple atomic analyses over a score to derive implicit musical evidence. To support diagnosis, we provide a taxonomy with six query intent categories and six analytical dimension tags. We further propose a tool-augmented retrieval and reasoning framework that integrates a ReAct-style controller with deterministic symbolic analysis operators built with music21. Experiments across prompting baselines and agent variants show that tool-grounded compositional retrieval consistently outperforms Large Language Model-only approaches, yielding 5-7% absolute accuracy gains, with the largest improvements on analysis-heavy categories.

音乐信息检索符号推理大模型应用工具增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。