arXiv:2604.13057cs.CL2026-04

分析孟加拉语和英语银行应用评论,发现传统模型比大模型更准。

A Multi-Model Approach to English-Bangla Sentiment Classification of Government Mobile Banking App Reviews

论文配图:A Multi-Model Approach to English-Bangla Sentiment Classification of Government Mobile Banking App Reviews
图 1 · 摘自论文原文
  • 用评分+分类器混合标注,提升标注一致性
  • 随机森林准确率达81.5%,优于微调的XLM-RoBERTa
  • 揭示界面慢、设计差是主要不满点,建议优先本地化

在发展中国家,数百万用户依赖移动银行获取金融服务,应用质量直接影响金融可及性。本研究分析了来自4款孟加拉国政府银行应用的5,652条英文与孟加拉语谷歌商店评论(原始数据11,414条)。采用评分与独立XLM-RoBERTa分类器结合的混合标注方法,获得中等一致率(kappa = 0.459)。传统模型表现优于基于Transformer的模型:随机森林准确率最高(0.815),线性SVM加权F1得分最高(0.804),均高于微调XLM-RoBERTa(0.793)。McNemar检验显示所有经典模型显著优于预训练XLM-RoBERTa(p < 0.05),而与微调版本差异不显著。使用DeBERTa-v3进行细粒度情感分析发现,用户主要抱怨交易速度慢和界面设计差;eJanata应用获最低评分。据此提出三项政策建议:改善应用质量、以信任为核心的发布管理、推行孟加拉语优先的NLP。值得注意的是,孟加拉语与英语之间存在16.1个百分点的准确率差距,凸显低资源语言模型开发的紧迫性。

原文摘要 · Abstract (English)

For millions of users in developing economies who depend on mobile banking as their primary gateway to financial services, app quality directly shapes financial access. The study analyzed 5,652 Google Play reviews in English and Bangla (filtered from 11,414 raw reviews) for four Bangladeshi government banking apps. The authors used a hybrid labeling approach that combined use of the reviewer's star rating for each review along with a separate independent XLM-RoBERTa classifier to produce moderate inter-method agreement (kappa = 0.459). Traditional models outperformed transformer-based ones: Random Forest produced the highest accuracy (0.815), while Linear SVM produced the highest weighted F1 score (0.804); both were higher than the performance of fine-tuned XLM-RoBERTa (0.793). McNemar's test confirmed that all classical models were significantly superior to the off-the-shelf XLM-RoBERTa (p < 0.05), while differences with the fine-tuned variant were not statistically significant. DeBERTa-v3 was applied to analyze the sentiment at the aspect level across the reviews for the four apps; the reviewers expressed their dissatisfaction primarily with the speed of transactions and with the poor design of interfaces; eJanata app received the worst ratings from the reviewers across all apps. Three policy recommendations are made based on these findings - remediation of app quality, trust-centred release management, and Bangla-first NLP adoption - to assist state-owned banks in moving towards improving their digital services through data-driven methods. Notably, a 16.1-percentage-point accuracy gap between Bangla and English text highlights the need for low-resource language model development.

情感分析低资源语言移动银行多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。