arXiv:2409.05521cs.CLcs.AI2024-09

测试大模型音乐推理能力,发现识谱困难暴露其逻辑短板

Harmonic Reasoning in Large Language Models

  • 用音乐任务测试大模型的逻辑推理能力
  • 模型对音程判断准确,但识别和弦与调式表现差
  • 适合研究模型认知局限或音乐人工智能方向

大型语言模型(LLMs)在艺术创作等任务中应用广泛,但在逻辑推理和计数等特定任务上仍存在挑战。本文考察了 GPT-3.5 与 GPT-4o 在音乐任务中的表现,包括从音程推断音符、识别和弦与调式。结果表明,模型在音程判断任务中表现良好,但在复杂任务如和弦与调式识别上表现不佳。该研究揭示了当前大模型在结构性逻辑推理上的明显局限,并提出了一个自动生成的基准数据集,以支持未来在音乐认知与模型改进方面的研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are becoming very popular and are used for many different purposes, including creative tasks in the arts. However, these models sometimes have trouble with specific reasoning tasks, especially those that involve logical thinking and counting. This paper looks at how well LLMs understand and reason when dealing with musical tasks like figuring out notes from intervals and identifying chords and scales. We tested GPT-3.5 and GPT-4o to see how they handle these tasks. Our results show that while LLMs do well with note intervals, they struggle with more complicated tasks like recognizing chords and scales. This points out clear limits in current LLM abilities and shows where we need to make them better, which could help improve how they think and work in both artistic and other complex areas. We also provide an automatically generated benchmark data set for the described tasks.

音乐推理大模型能力逻辑认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。