arXiv:2510.04291cs.CLcs.LG2025-10

融合BERT与决策树,提升波斯语情感分析准确率

PABSA: Hybrid Framework for Persian Aspect-Based Sentiment Analysis

  • 用多语言BERT极性分数做特征,输入决策树分类器
  • 在Pars-ABSA数据集上达到93.34%准确率,超越现有基准
  • 构建波斯语同义词与实体词典,支持文本增强

情感分析是自然语言处理中的关键任务,可从用户评论中提取有意义的见解。然而,由于标注数据稀缺、预处理工具有限,以及高质量嵌入和特征提取方法不足,波斯语的情感分析仍具挑战性。为此,我们提出一种结合机器学习与深度学习的混合方法,用于波斯语方面级情感分析(ABSA)。具体而言,我们利用多语言BERT生成的极性分数作为额外特征,并将其融入决策树分类器,在Pars-ABSA数据集上实现93.34%的准确率,超越现有基准。此外,我们构建了一个波斯语同义词与实体词典,作为新语言资源,支持通过同义词和命名实体替换进行文本增强。结果表明,混合建模与特征增强能有效推动低资源语言如波斯语的情感分析发展。

原文摘要 · Abstract (English)

Sentiment analysis is a key task in Natural Language Processing (NLP), enabling the extraction of meaningful insights from user opinions across various domains. However, performing sentiment analysis in Persian remains challenging due to the scarcity of labeled datasets, limited preprocessing tools, and the lack of high-quality embeddings and feature extraction methods. To address these limitations, we propose a hybrid approach that integrates machine learning (ML) and deep learning (DL) techniques for Persian aspect-based sentiment analysis (ABSA). In particular, we utilize polarity scores from multilingual BERT as additional features and incorporate them into a decision tree classifier, achieving an accuracy of 93.34%-surpassing existing benchmarks on the Pars-ABSA dataset. Additionally, we introduce a Persian synonym and entity dictionary, a novel linguistic resource that supports text augmentation through synonym and named entity replacement. Our results demonstrate the effectiveness of hybrid modeling and feature augmentation in advancing sentiment analysis for low-resource languages such as Persian.

情感分析波斯语混合模型文本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。