arXiv:2509.05617cs.CLcs.AI2025-09被引 2

构建歌词情感评估基准,测试大模型识别六种情绪强度的能力

From Joy to Fear: A Benchmark of Emotion Estimation in Pop Song Lyrics

  • 用多人评分法构建可靠歌词情感标注数据集
  • 零样本与微调模型在六类情绪识别上表现差异显著
  • 适合音乐情感分析与创意文本理解的研究者参考

歌曲歌词的情感内容对听众体验和音乐偏好具有关键影响。本文研究了多标签歌词情感归属任务,通过预测六种基本情绪的强度得分来实现。采用平均意见分(MOS)方法构建人工标注数据集,聚合多位人类评估者的标注以确保真实标签的可靠性。基于该数据集,我们在零样本场景下对多个公开的大语言模型(LLMs)进行了全面评估,并对一个基于BERT的模型进行微调以预测多标签情绪分数。实验结果揭示了零样本与微调模型在捕捉歌词情感细微差别的相对优劣。研究结果表明,大语言模型在创意文本情感识别中具有潜力,为基于情感的音乐信息检索应用提供了模型选择策略。标注数据集已公开:https://github.com/LLM-HITCS25S/LyricsEmotionAttribution。

原文摘要 · Abstract (English)

The emotional content of song lyrics plays a pivotal role in shaping listener experiences and influencing musical preferences. This paper investigates the task of multi-label emotional attribution of song lyrics by predicting six emotional intensity scores corresponding to six fundamental emotions. A manually labeled dataset is constructed using a mean opinion score (MOS) approach, which aggregates annotations from multiple human raters to ensure reliable ground-truth labels. Leveraging this dataset, we conduct a comprehensive evaluation of several publicly available large language models (LLMs) under zero-shot scenarios. Additionally, we fine-tune a BERT-based model specifically for predicting multi-label emotion scores. Experimental results reveal the relative strengths and limitations of zero-shot and fine-tuned models in capturing the nuanced emotional content of lyrics. Our findings highlight the potential of LLMs for emotion recognition in creative texts, providing insights into model selection strategies for emotion-based music information retrieval applications. The labeled dataset is available at https://github.com/LLM-HITCS25S/LyricsEmotionAttribution.

情感分析歌词理解大模型音乐信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。