让大模型精准定位音频中的音乐片段,提升语音理解的准确性。
What Are You Listening to? Temporal Music Grounding for Audio-to-Text Large Language Models

- 设计时间音乐对齐任务,让模型输出音乐事件的时间区间。
- 构建精确对齐的MGBench数据集,验证模型在三音和两小节片段中的定位能力。
- 发现专门训练可显著提升定位效果,适合音乐智能研究者使用。
大型音频-语言模型虽能生成流畅且符合音乐逻辑的回应,但其是否真正基于音频输入仍不明确。本文提出时间音乐对齐任务,要求模型返回与查询音乐音符、事件或模式对应的时间段。为此,我们构建了MusicGroundingBench基准测试集,通过算法生成钢琴MIDI并转为音频,实现符号与音频的精确对齐。该测试集包含两个子集:MGBench-3N用于评估最多含三个音符片段的音符级对齐;MGBench-2B用于评估两小节片段中的结构化对齐与短时音乐理解。实验表明,当前音频-语言模型在时间音乐对齐任务上仍面临挑战,而针对任务进行专项训练可带来显著性能提升。我们还初步探讨了对齐监督与音乐理解之间的关系。结果表明,MusicGroundingBench为评估模型是否基于时序音乐证据生成响应提供了可控的测试平台。
原文摘要 · Abstract (English)
Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce temporal music grounding, a task in which a model returns one or more time spans corresponding to a queried musical note, event, or pattern. To evaluate this capability, we present MusicGroundingBench, a controlled benchmark suite built by rendering algorithmically generated piano MIDI to audio, yielding exact symbolic-to-audio alignment. The suite comprises two subsets: MGBench-3N, which evaluates note-level grounding in clips containing up to three notes, and MGBench-2B, which evaluates structured grounding and short-form music understanding in two-bar excerpts. Experiments show that temporal music grounding remains challenging for current audio-language models, whereas task-specific training yields substantial gains. We further report exploratory evidence on the relationship between grounding supervision and music understanding. These results establish MusicGroundingBench as a controlled testbed for assessing whether audio-language models ground their responses in temporally localized musical evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。