arXiv:2605.29300cs.CLcs.AI2026-05被引 3

评测音乐大模型的时空定位能力,提出新方法显著提升精准度。

MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs

论文配图:MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs
图 1 · 摘自论文原文
  • 设计五类时序问答任务,用专家验证评估模型定位能力
  • 现有音乐大模型在时序定位上表现差,平均准确率不足40%
  • 提出四阶段优化方案,使定位准确率提升超25个百分点

近期大型音频-语言模型(LALMs)在理解音乐内容方面展现出良好潜力,但其回答是否准确对应音频中的时间区域仍缺乏深入研究。这一局限对音乐理解尤为关键,因关键信息常表现为时间局部事件,如乐器出现和节奏转换。为填补此空白,我们提出MusTBench,一个由音乐专家验证的基准,通过五类时序问答任务评估LALMs的时序定位能力。为进一步提升现有模型的时序定位能力,我们提出MusT,一种包含四个阶段的时序优化方案:音乐编码器适配、LLM适配、LLM监督微调以及基于强化学习的优化。在MusTBench上的实验表明,现有LALMs在精确时序定位上表现不佳,而MusT相比强基线有显著提升。这些结果确立了时序定位是当前LALMs的关键缺失能力,并将MusTBench定位为未来时序音乐理解研究的挑战性基准。

原文摘要 · Abstract (English)

Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temporal regions of the audio remains underexplored. This limitation is particularly critical for music understanding, where key information often occurs as temporally localized events, such as instrument entries and rhythmic transitions. To address this gap, we introduce MusTBench, a music-expert-validated benchmark designed to evaluate temporal grounding in LALMs through five temporally grounded question-answering tasks. To further improve temporal grounding in existing models, we propose MusT, a novel four-stage temporal optimization recipe spanning music encoder adaptation, LLM adaptation, LLM supervised fine-tuning, and RL-based optimization. Experiments on MusTBench show that existing LALMs struggle with precise temporal grounding, while MusT brings significant improvements over strong baselines. These results establish temporal grounding as a key missing capability in current LALMs and position MusTBench as a challenging benchmark for future research in temporally grounded music understanding.

音乐理解时序定位大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。