arXiv:2508.12626cs.SD2025-08中稿 · be published at IS…被引 1

用GPT-4o自动标注音乐情绪,效率高但细节略逊于人类。

Exploring the Feasibility of LLMs for Automated Music Emotion Annotation

  • 用GPT-4o对古典钢琴曲进行四象限情绪标注
  • 模型标注一致性在专家间差异范围内,整体准确率低于人工
  • 适合需要大规模标注的音乐情绪研究者使用

当前音乐情绪标注仍高度依赖人工,成本高昂,限制了数据规模。本研究评估了GPT-4o在音乐情绪标注中的可行性与可靠性。基于经典MIDI数据集GiantMIDI-Piano,采用四象限效价-唤醒框架,将GPT-4o的标注结果与三位人类专家对比。通过标准准确率、加权准确率(考虑专家间一致性)、标注者间一致性指标及标签分布相似性等多维度评估发现:尽管GPT-4o整体准确率低于人类专家,且在特定情绪状态分类上缺乏细腻度,但其标注变异范围处于专家自然差异区间内。结果表明,尽管存在局限,但基于GPT的标注在成本与效率上的优势,使其成为音乐情绪标注的有前景可扩展替代方案。

原文摘要 · Abstract (English)

Current approaches to music emotion annotation remain heavily reliant on manual labelling, a process that imposes significant resource and labour burdens, severely limiting the scale of available annotated data. This study examines the feasibility and reliability of employing a large language model (GPT-4o) for music emotion annotation. In this study, we annotated GiantMIDI-Piano, a classical MIDI piano music dataset, in a four-quadrant valence-arousal framework using GPT-4o, and compared against annotations provided by three human experts. We conducted extensive evaluations to assess the performance and reliability of GPT-generated music emotion annotations, including standard accuracy, weighted accuracy that accounts for inter-expert agreement, inter-annotator agreement metrics, and distributional similarity of the generated labels. While GPT's annotation performance fell short of human experts in overall accuracy and exhibited less nuance in categorizing specific emotional states, inter-rater reliability metrics indicate that GPT's variability remains within the range of natural disagreement among experts. These findings underscore both the limitations and potential of GPT-based annotation: despite its current shortcomings relative to human performance, its cost-effectiveness and efficiency render it a promising scalable alternative for music emotion annotation.

音乐情感大模型自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。