arXiv:2512.15547cs.CLcs.CY2025-12中稿 · 2025 28th Internat…被引 1

用机器学习分析孟加拉国2024年起义期间民众情绪,首次构建本土化语料库。

When a Nation Speaks: Machine Learning and NLP in People's Sentiment Analysis During Bangladesh's 2024 Mass Uprising

  • 基于2028条孟加拉语新闻标题,构建情感分类数据集
  • 语言专用模型优于多语言模型(准确率71%)
  • 揭示网络封锁等事件如何影响公众情绪波动

情感分析是自然语言处理的新兴领域,以往研究多聚焦于选举与社交媒体趋势,但在社会动荡中的情绪动态,尤其是孟加拉语语境下仍存在显著空白。本研究首次在国家危机背景下开展孟加拉语情感分析,通过收集2,028条来自主要脸书新闻门户的标注新闻标题,将其划分为愤怒、希望与绝望三类。利用隐狄利克雷分布(LDA)识别出政治腐败与公众抗议等核心主题,并分析互联网封锁等事件对情绪模式的影响。所提方法在准确率上超越多语言模型(mBERT: 67%,XLM-RoBERTa: 71%)及传统机器学习方法(SVM与逻辑回归:均为70%)。结果表明语言特异性模型在危机情绪分析中更具优势,为理解政治动荡时期公众心理提供了新视角。

原文摘要 · Abstract (English)

Sentiment analysis, an emerging research area within natural language processing (NLP), has primarily been explored in contexts like elections and social media trends, but there remains a significant gap in understanding emotional dynamics during civil unrest, particularly in the Bangla language. Our study pioneers sentiment analysis in Bangla during a national crisis by examining public emotions amid Bangladesh's 2024 mass uprising. We curated a unique dataset of 2,028 annotated news headlines from major Facebook news portals, classifying them into Outrage, Hope, and Despair. Through Latent Dirichlet Allocation (LDA), we identified prevalent themes like political corruption and public protests, and analyzed how events such as internet blackouts shaped sentiment patterns. It outperformed multilingual transformers (mBERT: 67%, XLM-RoBERTa: 71%) and traditional machine learning methods (SVM and Logistic Regression: both 70%). These results highlight the effectiveness of language-specific models and offer valuable insights into public sentiment during political turmoil.

情感分析危机舆情孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。