arXiv:2501.13261cs.IRcs.SD2025-01被引 3

用GPT检测音乐信息检索中的标注错误,效果优于随机猜测。

Exploring GPT's Ability as a Judge in Music Understanding

  • 将音乐数据转为符号输入,通过提示工程让GPT判断任务标注是否正确。
  • 在节拍、和弦、调性任务中准确率分别达65.20%、64.80%、59.72%。
  • 提供更多音乐概念信息能提升GPT的判断一致性,适合音乐与AI交叉研究者。

近期基于文本的大语言模型(LLM)及其多模态感知能力的发展,促使我们探索其在音乐信息检索(MIR)挑战中的应用。本文采用系统化的提示工程方法,将音乐数据转换为符号输入,评估LLM在三大关键MIR任务中的标注错误检测能力:节拍跟踪、和弦提取和调性估计。提出一种概念增强方法,以检验LLM在提示中提供音乐概念时的推理一致性。实验测试了生成式预训练变换器(GPT)的MIR能力。结果表明,GPT在节拍跟踪、和弦提取和调性估计任务中的错误检测准确率分别为65.20%、64.80%和59.72%,均高于随机基线。此外,观察到GPT的错误发现准确率与提示中提供的概念信息量呈正相关。这些基于符号音乐输入的发现为未来基于LLM的MIR研究奠定了坚实基础。

原文摘要 · Abstract (English)

Recent progress in text-based Large Language Models (LLMs) and their extended ability to process multi-modal sensory data have led us to explore their applicability in addressing music information retrieval (MIR) challenges. In this paper, we use a systematic prompt engineering approach for LLMs to solve MIR problems. We convert the music data to symbolic inputs and evaluate LLMs' ability in detecting annotation errors in three key MIR tasks: beat tracking, chord extraction, and key estimation. A concept augmentation method is proposed to evaluate LLMs' music reasoning consistency with the provided music concepts in the prompts. Our experiments tested the MIR capabilities of Generative Pre-trained Transformers (GPT). Results show that GPT has an error detection accuracy of 65.20%, 64.80%, and 59.72% in beat tracking, chord extraction, and key estimation tasks, respectively, all exceeding the random baseline. Moreover, we observe a positive correlation between GPT's error finding accuracy and the amount of concept information provided. The current findings based on symbolic music input provide a solid ground for future LLM-based MIR research.

音乐理解大模型错误检测提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。