arXiv:2511.18146cs.CL2025-11被引 2

构建首个斯里兰卡僧伽罗语音乐评论情感数据集,支持情绪分析研究。

GeeSanBhava: Sentiment Tagged Sinhala Music Video Comment Data Set

  • 人工标注评论情感,采用情感维度模型确保标注一致性。
  • 多层感知机模型经调参后达到0.887的ROC-AUC,性能良好。
  • 数据集可助力僧伽罗语自然语言处理与音乐情绪识别研究。

本研究提出GeeSanBhava,一个高质量的僧伽罗语歌曲评论数据集,从YouTube手动采集并由三位独立标注者使用Russell情感维度模型(效价-唤醒度)进行标注,标注者间一致性达Fleiss kappa = 84.96%。分析显示不同歌曲具有明显的情感特征,凸显基于评论的情绪映射价值。研究还探讨了评论情感与歌曲情感的对比挑战,缓解用户生成内容中的固有偏见。利用相关大规模僧伽罗语新闻评论数据集预训练的多种机器学习与深度学习模型,报告了该数据集的零样本结果。经过超参数优化的三层多层感知机(256, 128, 64神经元)取得0.887的ROC-AUC分数。本研究为僧伽罗语自然语言处理和音乐情感识别提供了重要资源与洞见。

原文摘要 · Abstract (English)

This study introduce GeeSanBhava, a high-quality data set of Sinhala song comments extracted from YouTube manually tagged using Russells Valence-Arousal model by three independent human annotators. The human annotators achieve a substantial inter-annotator agreement (Fleiss kappa = 84.96%). The analysis revealed distinct emotional profiles for different songs, highlighting the importance of comment based emotion mapping. The study also addressed the challenges of comparing comment-based and song-based emotions, mitigating biases inherent in user-generated content. A number of Machine learning and deep learning models were pre-trained on a related large data set of Sinhala News comments in order to report the zero-shot result of our Sinhala YouTube comment data set. An optimized Multi-Layer Perceptron model, after extensive hyperparameter tuning, achieved a ROC-AUC score of 0.887. The model is a three-layer MLP with a configuration of 256, 128, and 64 neurons. This research contributes a valuable annotated dataset and provides insights for future work in Sinhala Natural Language Processing and music emotion recognition.

情感分析语音数据集僧伽罗语音乐情绪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。