arXiv:2410.05778cs.CLcs.LG2024-10

用Reddit评论训练模型,解决歌词情感分类数据少难题

Song Emotion Classification of Lyrics with Out-of-Domain Data under Label Scarcity

  • 用大规模外部数据(Reddit评论)替代稀缺的歌词数据
  • CNN模型在歌词情感分类任务上表现良好,准确率可观
  • 适合数据匮乏领域的情感分析研究者参考

歌曲对人类情绪有深远影响,歌词尤其能激发听众情绪变化。然而,高质量、大规模的歌词情感分类领域内数据集稀缺(Edmonds and Sedoc, 2021;Zhou, 2022)。领域内训练数据难以获取(Zhang and Miao, 2023),且标注成本高、耗时长(Azad et al., 2018)。本文探索利用大规模外部数据作为创新解决方案,应对歌词情感分类中训练数据不足的问题。实验发现,基于大型Reddit评论数据集训练的CNN模型,在歌词情感分类任务上表现出令人满意的性能与泛化能力,为缺乏领域内数据或获取成本高的场景提供了可行路径。

原文摘要 · Abstract (English)

Songs have been found to profoundly impact human emotions, with lyrics having significant power to stimulate emotional changes in the audience. There is a scarcity of large, high quality in-domain datasets for lyrics-based song emotion classification (Edmonds and Sedoc, 2021; Zhou, 2022). It has been noted that in-domain training datasets are often difficult to acquire (Zhang and Miao, 2023) and that label acquisition is often limited by cost, time, and other factors (Azad et al., 2018). We examine the novel usage of a large out-of-domain dataset as a creative solution to the challenge of training data scarcity in the emotional classification of song lyrics. We find that CNN models trained on a large Reddit comments dataset achieve satisfactory performance and generalizability to lyrical emotion classification, thus giving insights into and a promising possibility in leveraging large, publicly available out-of-domain datasets for domains whose in-domain data are lacking or costly to acquire.

情感分类小样本跨域训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。