BASS评测音频大模型在音乐结构与语义推理上的表现,揭示当前模型短板。
BASS: Benchmarking Audio LMs for Musical Structure and Semantic Reasoning
- 构建跨4类任务的音乐理解评测集,涵盖结构分割等12项任务
- 14个模型在结构分割和合作分析上表现差,歌词转录最擅长
- 适合音乐信息检索、推荐系统研究者参考
音乐理解需综合判断音频的结构与语义特征。我们提出BASS,用于评估音频语言模型在四大类别上的音乐理解与推理能力:结构分割、歌词转录、音乐学分析及艺术家合作。BASS包含2658个问题,覆盖12项任务、1993首独特歌曲,时长超138小时,涵盖多种流派,旨在检验真实场景下的音乐学知识与推理能力。我们评估了14个开源及前沿多模态大模型,发现即使最先进的模型在结构分割和艺术家合作等高层推理任务上仍表现不佳,而歌词转录表现最优。分析表明,当前模型能有效利用语言先验,但在音乐结构、人声与音乐学属性推理方面仍受限。BASS为音乐推荐与搜索等应用提供评估框架,有助于推动音频大模型发展。
原文摘要 · Abstract (English)
Music understanding is a complex task that often requires reasoning over both structural and semantic elements of audio. We introduce BASS, designed to evaluate music understanding and reasoning in audio language models across four broad categories: structural segmentation, lyric transcription, musicological analysis, and artist collaboration. BASS comprises 2658 questions spanning 12 tasks, 1993 unique songs and covering over 138 hours of music from a wide range of genres and tracks, crafted to assess musicological knowledge and reasoning in real-world scenarios. We evaluate 14 open-source and frontier multimodal LMs, finding that even state-of-the-art models struggle on higher-level reasoning tasks such as structural segmentation and artist collaboration, while performing best on lyric transcription. Our analysis reveals that current models leverage linguistic priors effectively but remain limited in reasoning over musical structure, vocal, and musicological attributes. BASS provides an evaluation framework with widespread applications in music recommendation and search and has the potential to guide the development of audio LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。