针对孟加拉语电商评论,构建可解释的跨领域情感分析框架
BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning

- 融合多种模型的动态加权集成学习方法
- 准确率85%,F1值0.88,跨域零样本仍保持67%-76%效果
- 支持可解释性分析,适合低资源语言场景应用
由于标注数据有限、形态复杂、混合编码现象及领域偏移问题,孟加拉语电商评论的多方面情感分析对3亿孟加拉语用户构成挑战。现有方法缺乏可解释性与跨领域泛化能力。本文提出BanglaSentNet,一种结合LSTM、BiLSTM、GRU与BanglaBERT的可解释混合深度学习框架,采用动态加权集成学习进行多方面情感分类。构建了包含8,755条人工标注的孟加拉语产品评论数据集,涵盖质量、服务、价格、装饰四个维度。引入基于SHAP的特征归因与注意力可视化以实现透明洞察。实验表明,该框架在准确率上达85%,F1值为0.88,比单一模型高出3-7%,显著优于传统方法。可解释性模块获得9.4/10的评分,人类一致性达87.6%。跨领域迁移学习显示:零样本下在不同领域(书籍评论、社交媒体、通用电商、新闻标题)仍保持67%-76%有效性;少量样本(500-1000)微调即可达到全量微调90%-95%的效果,大幅降低标注成本。实际部署验证其在孟加拉电商平台中的实用性,助力定价优化、服务改进与用户体验提升。本研究建立了孟加拉语情感分析新基准,推动低资源语言集成学习发展,并提供商业可行方案。
原文摘要 · Abstract (English)
Multi-aspect sentiment analysis of Bangla e-commerce reviews remains challenging due to limited annotated datasets, morphological complexity, code-mixing phenomena, and domain shift issues, affecting 300 million Bangla-speaking users. Existing approaches lack explainability and cross-domain generalization capabilities crucial for practical deployment. We present BanglaSentNet, an explainable hybrid deep learning framework integrating LSTM, BiLSTM, GRU, and BanglaBERT through dynamic weighted ensemble learning for multi-aspect sentiment classification. We introduce a dataset of 8,755 manually annotated Bangla product reviews across four aspects (Quality, Service, Price, Decoration) from major Bangladeshi e-commerce platforms. Our framework incorporates SHAP-based feature attribution and attention visualization for transparent insights. BanglaSentNet achieves 85% accuracy and 0.88 F1-score, outperforming standalone deep learning models by 3-7% and traditional approaches substantially. The explainability suite achieves 9.4/10 interpretability score with 87.6% human agreement. Cross-domain transfer learning experiments reveal robust generalization: zero-shot performance retains 67-76% effectiveness across diverse domains (BanglaBook reviews, social media, general e-commerce, news headlines); few-shot learning with 500-1000 samples achieves 90-95% of full fine-tuning performance, significantly reducing annotation costs. Real-world deployment demonstrates practical utility for Bangladeshi e-commerce platforms, enabling data-driven decision-making for pricing optimization, service improvement, and customer experience enhancement. This research establishes a new state-of-the-art benchmark for Bangla sentiment analysis, advances ensemble learning methodologies for low-resource languages, and provides actionable solutions for commercial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。