arXiv:2411.19584cs.LG2024-11被引 11

融合词典与预训练模型,提升孟加拉语情感分析精度

Enhancing Sentiment Analysis in Bengali Texts: A Hybrid Approach Using Lexicon-Based Algorithm and Pretrained Language Model Bangla-BERT

  • 构建规则算法BSPS,结合词典打分生成情感得分
  • 在1.5万条标注数据上,九分类准确率优于纯模型方法
  • 适合资源匮乏语言的情感分析研究者参考

情感分析旨在识别文本中的情绪倾向与极性,揭示用户复杂情感。尽管英语等语言的分析已较成熟,孟加拉语的研究仍有限,尤其在细粒度分类方面。本文提出一种新方法,融合基于规则的算法与预训练语言模型。我们从零构建了超过1.5万条人工标注的评论数据集,并建立词典数据表,为评论分配情感极性分值。提出新型规则算法Bangla Sentiment Polarity Score(BSPS),可生成情感得分并实现九类情感分类。通过使用预训练的孟加拉语Transformer模型BanglaBERT评估分类结果,并对比其在原始数据上的直接分类表现。实验表明,BSPS+BanglaBERT混合方法在九类情感分类中均取得更高准确率、精确率,显著优于单独使用BanglaBERT。研究验证了规则方法与预训练模型结合在孟加拉语情感分析中的有效性,为类似语言研究提供新路径。

原文摘要 · Abstract (English)

Sentiment analysis (SA) is a process of identifying the emotional tone or polarity within a given text and aims to uncover the user's complex emotions and inner feelings. While sentiment analysis has been extensively studied for languages like English, research in Bengali, remains limited, particularly for fine-grained sentiment categorization. This work aims to connect this gap by developing a novel approach that integrates rule-based algorithms with pre-trained language models. We developed a dataset from scratch, comprising over 15,000 manually labeled reviews. Next, we constructed a Lexicon Data Dictionary, assigning polarity scores to the reviews. We developed a novel rule based algorithm Bangla Sentiment Polarity Score (BSPS), an approach capable of generating sentiment scores and classifying reviews into nine distinct sentiment categories. To assess the performance of this method, we evaluated the classified sentiments using BanglaBERT, a pre-trained transformer-based language model. We also performed sentiment classification directly with BanglaBERT on the original data and evaluated this model's results. Our analysis revealed that the BSPS + BanglaBERT hybrid approach outperformed the standalone BanglaBERT model, achieving higher accuracy, precision, and nuanced classification across the nine sentiment categories. The results of our study emphasize the value and effectiveness of combining rule-based and pre-trained language model approaches for enhanced sentiment analysis in Bengali and suggest pathways for future research and application in languages with similar linguistic complexities.

情感分析孟加拉语混合方法预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。