arXiv:2507.11084cs.CL2025-07中稿 · and presented at t…被引 11

用混合Transformer模型分析孟加拉语社交媒体情绪,准确率达83.7%。

Social Media Sentiments Analysis on the July Revolution in Bangladesh: A Hybrid Transformer Based Machine Learning Approach

  • 融合多种BERT模型与投票机制,提升低资源语言情感识别能力
  • 在4200条孟加拉语评论上实现83.7%的准确率,优于单一模型
  • 适合关注社会运动分析、低资源语言NLP的研究者

孟加拉国7月革命是一场由学生主导的大规模抗议活动,民众通过社交媒体表达对正义与制度改革的诉求。本文构建了一个基于混合Transformer的情感分析框架,用于解析革命期间及之后社交媒体评论中的公众情绪。研究使用全新收集的4,200条孟加拉语评论数据集,采用BanglaBERT、mBERT、XLM-RoBERTa及自研的XMB-BERT等多模型进行特征提取,并通过主成分分析(PCA)降低维度以提高计算效率。对比了11种传统与先进机器学习分类器,最终提出的XMB-BERT结合投票分类器取得83.7%的准确率,显著优于其他组合。该研究展示了机器学习在低资源语言如孟加拉语中分析社会情绪的潜力。

原文摘要 · Abstract (English)

The July Revolution in Bangladesh marked a significant student-led mass uprising, uniting people across the nation to demand justice, accountability, and systemic reform. Social media platforms played a pivotal role in amplifying public sentiment and shaping discourse during this historic mass uprising. In this study, we present a hybrid transformer-based sentiment analysis framework to decode public opinion expressed in social media comments during and after the revolution. We used a brand new dataset of 4,200 Bangla comments collected from social media. The framework employs advanced transformer-based feature extraction techniques, including BanglaBERT, mBERT, XLM-RoBERTa, and the proposed hybrid XMB-BERT, to capture nuanced patterns in textual data. Principle Component Analysis (PCA) were utilized for dimensionality reduction to enhance computational efficiency. We explored eleven traditional and advanced machine learning classifiers for identifying sentiments. The proposed hybrid XMB-BERT with the voting classifier achieved an exceptional accuracy of 83.7% and outperform other model classifier combinations. This study underscores the potential of machine learning techniques to analyze social sentiment in low-resource languages like Bangla.

情感分析低资源语言Transformer社会运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。