构建尼泊尔语社交媒体情感多标签分析数据集,揭示情绪演化规律
NepEMO: A Multi-Label Emotion and Sentiment Analysis on Nepali Reddit with Linguistic Insights and Temporal Trends
- 构建4462条尼泊尔语帖子的多标签情感标注数据集
- 变压器模型在情感与情绪分类任务中表现最佳
- 揭示情绪共现模式与语言特征,适用于南亚语种研究
社交媒体平台(如Facebook、Twitter和Reddit)在自然灾害、疫情、选举等重大事件期间,成为人们表达意见与情绪的重要渠道。其中,Reddit因支持匿名发言,特别适合讨论健康、生活等敏感话题。本文提出一个新数据集NepEMO,用于尼泊尔语子版块的多标签情绪(MLE)与情感分类(SC)。该数据集包含4462条手动标注的帖子,涵盖英语、罗马化尼泊尔语和天城文三种书写形式,覆盖五类情绪(恐惧、愤怒、悲伤、喜悦、抑郁)与三类情感(正向、负向、中性),时间跨度为2019年1月至2025年6月。我们通过主题建模(LDA)、TF-IDF关键词提取等方法分析语言特征,包括情绪趋势、情绪共现模式与情感特异性n-gram。此外,对比了多种传统机器学习、深度学习与变换器模型。结果表明,变换器模型在两项任务中均显著优于其他模型。
原文摘要 · Abstract (English)
Social media (SM) platforms (e.g. Facebook, Twitter, and Reddit) are increasingly leveraged to share opinions and emotions, specifically during challenging events, such as natural disasters, pandemics, and political elections, and joyful occasions like festivals and celebrations. Among the SM platforms, Reddit provides a unique space for its users to anonymously express their experiences and thoughts on sensitive issues such as health and daily life. In this work, we present a novel dataset, called NepEMO, for multi-label emotion (MLE) and sentiment classification (SC) on the Nepali subreddit post. We curate and build a manually annotated dataset of 4,462 posts (January 2019- June 2025) written in English, Romanised Nepali and Devanagari script for five emotions (fear, anger, sadness, joy, and depression) and three sentiment classes (positive, negative, and neutral). We perform a detailed analysis of posts to capture linguistic insights, including emotion trends, co-occurrence of emotions, sentiment-specific n-grams, and topic modelling using Latent Dirichlet Allocation and TF-IDF keyword extraction. Finally, we compare various traditional machine learning (ML), deep learning (DL), and transformer models for MLE and SC tasks. The result shows that transformer models consistently outperform the ML and DL models for both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。