构建音乐指令遵循评测基准,推动音频-文本大模型发展
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
- 将传统音乐信息检索任务转为指令跟随格式,覆盖20余项核心任务
- 实测显示大模型性能显著低于专业模型,存在文化与性别偏差
- 支持多款开源模型评测,为音乐大模型提供统一评估框架
近年来,音频-文本大语言模型(LLMs)在音乐理解与生成方面取得进展,但现有评测基准范围有限,常依赖简化任务或选择题评估,难以反映真实音乐分析的复杂性。本文将多种传统音乐信息检索(MIR)标注重新诠释为指令跟随形式,提出CMI-Bench——一个涵盖风格分类、情绪回归、乐器识别、音高估计、调性检测、歌词转录、旋律提取、人声技巧识别、演奏技法检测、音乐标签、音乐描述和节拍追踪等20余项任务的综合性音乐指令遵循评测基准。该基准采用与先进监督模型一致的标准化评价指标,确保与传统方法可直接对比。我们提供了支持LTU、Qwen-audio、SALMONN、MusiLingo等开源音频-文本大模型的评测工具包。实验表明,当前大模型在多数任务上表现远逊于监督模型,且存在文化、时间与性别偏见,揭示其在处理音乐任务中的潜力与局限。CMI-Bench为音乐指令遵循能力评估建立统一基础,推动音乐感知型大模型的发展。
原文摘要 · Abstract (English)
Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or multi-choice evaluations that fail to reflect the complexity of real-world music analysis. We reinterpret a broad range of traditional MIR annotations as instruction-following formats and introduce CMI-Bench, a comprehensive music instruction following benchmark designed to evaluate audio-text LLMs on a diverse set of music information retrieval (MIR) tasks. These include genre classification, emotion regression, emotion tagging, instrument classification, pitch estimation, key detection, lyrics transcription, melody extraction, vocal technique recognition, instrument performance technique detection, music tagging, music captioning, and (down)beat tracking: reflecting core challenges in MIR research. Unlike previous benchmarks, CMI-Bench adopts standardized evaluation metrics consistent with previous state-of-the-art MIR models, ensuring direct comparability with supervised approaches. We provide an evaluation toolkit supporting all open-source audio-textual LLMs, including LTU, Qwen-audio, SALMONN, MusiLingo, etc. Experiment results reveal significant performance gaps between LLMs and supervised models, along with their culture, chronological and gender bias, highlighting the potential and limitations of current models in addressing MIR tasks. CMI-Bench establishes a unified foundation for evaluating music instruction following, driving progress in music-aware LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。