arXiv:2502.01091cs.CLcs.AI2025-02被引 3

用ParsBERT和词典提升波斯语情感分析准确率

Enhancing Aspect-based Sentiment Analysis with ParsBERT in Persian Language

  • 基于ParsBERT模型结合领域词典进行细粒度情感分析
  • 在Digikala数据上达88.2%准确率与61.7%F1值
  • 适合做波斯语社交媒体情感分析的研究者参考

在互联网普及与社交网络主导的背景下,波斯语文本挖掘面临数据集匮乏和现有语言模型效率低下的挑战。本文针对这些问题,旨在提升专用于波斯语的语言模型效能。通过采用基于方面的分析方法,结合ParsBERT模型与相关词典,对来自波斯语网站Digikala的用户评论进行情感分析。实验结果表明,该方法不仅具备优越的语义理解能力,还显著提升了效率:准确率达88.2%,F1得分为61.7。提升语言模型在此领域的表现,对于从用户生成内容中提取细微情感至关重要,将推动波斯语情感分析在效率与准确性上的发展。

原文摘要 · Abstract (English)

In the era of pervasive internet use and the dominance of social networks, researchers face significant challenges in Persian text mining including the scarcity of adequate datasets in Persian and the inefficiency of existing language models. This paper specifically tackles these challenges, aiming to amplify the efficiency of language models tailored to the Persian language. Focusing on enhancing the effectiveness of sentiment analysis, our approach employs an aspect-based methodology utilizing the ParsBERT model, augmented with a relevant lexicon. The study centers on sentiment analysis of user opinions extracted from the Persian website 'Digikala.' The experimental results not only highlight the proposed method's superior semantic capabilities but also showcase its efficiency gains with an accuracy of 88.2% and an F1 score of 61.7. The importance of enhancing language models in this context lies in their pivotal role in extracting nuanced sentiments from user-generated content, ultimately advancing the field of sentiment analysis in Persian text mining by increasing efficiency and accuracy.

波斯语情感分析ParsBERT细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。