提出新评估方法,精准测试大模型对音乐事实的理解能力
Assessing Factual Music Comprehension in Large Audio Language Models
- 设计结构化评估协议,将开放回答转为可量化指标
- 在3个数据集上构建6项事实检索任务,覆盖多样音乐场景
- 公开工具链,支持新模型快速评测,适合音乐AI研究者
大型音频语言模型(LALMs)利用多模态表示生成针对音频的自然语言问答。本文发现,现有MusicQA数据集无法有效衡量模型回答的音乐事实正确性,并提出新的评估协议:通过提示模型输出可验证的事实信息,将开放式回答解析为结构化内容,使用精确率、召回率和F1值进行客观评估。基于此,我们在MusicNet、Free Music Archive和OverClocked ReMix三个多样化数据集上构建了六个事实信息检索任务,对九个近期LALMs(包括Gemini和Music Flamingo等前沿模型)进行了基准测试,并在https://github.com/DCL2004/LALM-Eval发布全套评估脚本,以促进新模型的评测。
原文摘要 · Abstract (English)
Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs using the popular MusicQA dataset fails to measure whether a model's responses about music are factually correct, and (2) develop a new protocol for assessing the music comprehension capabilities of LALMs. Specifically, we propose an evaluation protocol that prompts a LALM for factually verifiable information, and parses its open-ended response into a structured format that can be objectively assessed using Precision, Recall, and F1 scores. Using this protocol, we define a benchmark consisting of six factual information retrieval tasks defined on three diverse datasets: MusicNet, the Free Music Archive, and OverClocked ReMix. We benchmark nine recent LALMs, including frontier models like Gemini and the latest open models like Music Flamingo, and release the suite of evaluation scripts at https://github.com/DCL2004/LALM-Eval to facilitate benchmarking of new LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。