arXiv:2506.10154cs.CLcs.LG2025-06被引 5

用机器学习分析孟加拉语社交评论情绪,提升低资源语言情感识别能力

Analyzing Emotions in Bangla Social Media Comments Using Machine Learning and LIME

  • 采用SVM、KNN、随机森林等模型结合TF-IDF与n-gram特征
  • 基于22,698条数据的实验表明AdaBoost在情绪分类中表现最优
  • 使用LIME解释模型决策,增强可解释性,适合低资源语言研究者

针对缺乏资源的孟加拉语,本研究基于22,698条来自EmoNoBa数据集的社交媒体评论开展情绪分析。采用线性SVM、KNN和随机森林等机器学习模型,结合TF-IDF向量化后的n-gram特征进行语言建模,并探索主成分分析(PCA)对降维的影响。此外,引入双向长短期记忆网络(BiLSTM)与AdaBoost改进决策树性能。为提升模型可解释性,利用LIME对AdaBoost分类器的预测结果进行解释。研究旨在推动低资源语言中的情感分析技术发展,验证多种方法在孟加拉语情绪识别中的有效性。

原文摘要 · Abstract (English)

Research on understanding emotions in written language continues to expand, especially for understudied languages with distinctive regional expressions and cultural features, such as Bangla. This study examines emotion analysis using 22,698 social media comments from the EmoNoBa dataset. For language analysis, we employ machine learning models: Linear SVM, KNN, and Random Forest with n-gram data from a TF-IDF vectorizer. We additionally investigated how PCA affects the reduction of dimensionality. Moreover, we utilized a BiLSTM model and AdaBoost to improve decision trees. To make our machine learning models easier to understand, we used LIME to explain the predictions of the AdaBoost classifier, which uses decision trees. With the goal of advancing sentiment analysis in languages with limited resources, our work examines various techniques to find efficient techniques for emotion identification in Bangla.

情感分析孟加拉语可解释性机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。