评测发现大模型在音乐理解上易受提示影响,需针对性适配才能用。
Evaluation of pretrained language models on music understanding
- 用音类层级三元组测试模型对音乐标签相对关系的判断能力
- 六款模型均出现误判,尤其在否定句和关键词敏感性上表现差
- 适合研究音乐与文本融合任务的开发者参考模型局限性
音乐-文本多模态系统推动了音乐信息检索等应用的发展,但对大语言模型(LLM)音乐知识的评估仍不足。本文通过三元组准确性评估,揭示了LLM存在提示敏感、无法建模否定(如'无吉他摇滚乐')以及对特定词汇敏感等问题。基于Audioset层级本体生成包含锚点、正例与负例标签的三元组,评估了六种通用Transformer模型在流派与乐器子树上的表现。尽管整体准确率较高,但所有模型均显现出不一致性,表明现成的LLM需针对音乐任务进行适配后方可使用。
原文摘要 · Abstract (English)
Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the reported success, little effort has been put into evaluating the musical knowledge of Large Language Models (LLM). In this paper, we demonstrate that LLMs suffer from 1) prompt sensitivity, 2) inability to model negation (e.g. 'rock song without guitar'), and 3) sensitivity towards the presence of specific words. We quantified these properties as a triplet-based accuracy, evaluating the ability to model the relative similarity of labels in a hierarchical ontology. We leveraged the Audioset ontology to generate triplets consisting of an anchor, a positive (relevant) label, and a negative (less relevant) label for the genre and instruments sub-tree. We evaluated the triplet-based musical knowledge for six general-purpose Transformer-based models. The triplets obtained through this methodology required filtering, as some were difficult to judge and therefore relatively uninformative for evaluation purposes. Despite the relatively high accuracy reported, inconsistencies are evident in all six models, suggesting that off-the-shelf LLMs need adaptation to music before use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。