arXiv:2511.21381cs.LGcs.CL2025-11被引 2

首个孟加拉语电商评论三元组抽取框架,提升细粒度情感分析效果。

BanglaASTE: A Novel Framework for Aspect-Sentiment-Opinion Extraction in Bangla E-commerce Reviews Using Ensemble Deep Learning

  • 融合图结构与语义相似度的混合分类方法
  • 集成孟加拉语BERT与XGBoost模型,准确率达89.9%
  • 适合低资源语言情感分析研究者与电商领域应用

基于方面的情感分析(ABSA)已成为从用户生成内容中提取细粒度情感洞察的关键工具,尤其在电商与社交媒体领域。然而,由于缺乏全面的数据集和针对该语言的三元组抽取专用框架,孟加拉语ABSA研究仍严重不足。本文提出BanglaASTE,一个用于孟加拉语产品评论的方面-观点-情感三元组抽取新框架。贡献包括:(1) 构建首个标注的孟加拉语ASTE数据集,包含3,345条来自Daraz、Facebook、Rokomari等主流电商平台的评论;(2) 开发一种结合图结构方面-观点匹配与语义相似度技术的混合分类框架;(3) 实现集成孟加拉语BERT上下文嵌入与XGBoost增强算法的集成模型,显著提升三元组抽取性能。实验结果表明,该集成方法在准确率上达到89.9%,F1分数为89.1%,在各项评估指标上均显著优于基线模型。该框架有效应对了孟加拉语文本处理中的非正式表达、拼写变异和数据稀疏等关键挑战。本研究推动了低资源语言情感分析的前沿进展,并为孟加拉语电商分析提供了可扩展解决方案。

原文摘要 · Abstract (English)

Aspect-Based Sentiment Analysis (ABSA) has emerged as a critical tool for extracting fine-grained sentiment insights from user-generated content, particularly in e-commerce and social media domains. However, research on Bangla ABSA remains significantly underexplored due to the absence of comprehensive datasets and specialized frameworks for triplet extraction in this language. This paper introduces BanglaASTE, a novel framework for Aspect Sentiment Triplet Extraction (ASTE) that simultaneously identifies aspect terms, opinion expressions, and sentiment polarities from Bangla product reviews. Our contributions include: (1) creation of the first annotated Bangla ASTE dataset containing 3,345 product reviews collected from major e-commerce platforms including Daraz, Facebook, and Rokomari; (2) development of a hybrid classification framework that employs graph-based aspect-opinion matching with semantic similarity techniques; and (3) implementation of an ensemble model combining BanglaBERT contextual embeddings with XGBoost boosting algorithms for enhanced triplet extraction performance. Experimental results demonstrate that our ensemble approach achieves superior performance with 89.9% accuracy and 89.1% F1-score, significantly outperforming baseline models across all evaluation metrics. The framework effectively addresses key challenges in Bangla text processing including informal expressions, spelling variations, and data sparsity. This research advances the state-of-the-art in low-resource language sentiment analysis and provides a scalable solution for Bangla e-commerce analytics applications.

情感分析孟加拉语三元组抽取深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。