测试大模型在音乐学中的可靠性,发现加知识库的模型更准
The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
- 用检索增强生成和多选题自动生成基准测试
- 400个专家验证题中,带知识库的模型表现更优
- 适合音乐学与AI交叉研究者参考
本文探讨大型语言模型(LLMs)在音乐学中的应用与可靠性。通过与专家和学生的讨论,评估了当前对这一普及技术的接受度与担忧。我们进一步提出一种半自动方法,利用检索增强生成模型和多选题生成构建初始基准,并由人类专家验证。在400个经人工验证的问题上评估显示,当前的通用大模型可靠性低于基于音乐词典的检索增强生成模型。论文建议,要释放大模型在音乐学中的潜力,需开展以音乐学为导向的研究,通过引入准确可靠的领域知识来定制专用大模型。
原文摘要 · Abstract (English)
In this work, we explore the use and reliability of Large Language Models (LLMs) in musicology. From a discussion with experts and students, we assess the current acceptance and concerns regarding this, nowadays ubiquitous, technology. We aim to go one step further, proposing a semi-automatic method to create an initial benchmark using retrieval-augmented generation models and multiple-choice question generation, validated by human experts. Our evaluation on 400 human-validated questions shows that current vanilla LLMs are less reliable than retrieval augmented generation from music dictionaries. This paper suggests that the potential of LLMs in musicology requires musicology driven research that can specialized LLMs by including accurate and reliable domain knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。