arXiv:2608.01471cs.CL2026-08

用持续学习提升孟加拉语情感分类,兼顾效率与可解释性

Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning

论文配图:Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning
图 1 · 摘自论文原文
  • 分两阶段:先领域自适应预训练,再参数高效微调
  • 在新闻数据上性能稳定,接近强基线模型
  • 结合SHAP分析语法特征对情感判断的影响

低资源语言的情感理解仍是自然语言处理的关键挑战,尤其在领域特定数据稀缺时。本文提出 SentiBanglaBERT,一种结合领域自适应持续预训练与参数高效微调的两阶段孟加拉语情感分类框架。该方法在保持计算高效的同时,实现对新闻类文本的上下文适应,采用低秩适配(LoRA)技术。实验表明,其性能稳定且接近强基线模型,同时集成基于SHAP的可解释性分析,揭示了否定后缀、体标记等孟加拉语形态特征如何影响情感预测。该框架展示了领域自适应持续学习在形态丰富、资源匮乏语言中的潜力,为高效、可解释的NLP提供新路径。

原文摘要 · Abstract (English)

Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali sentiment classification framework combining domain-adaptive continual pretraining and parameter-efficient fine-tuning. The approach enables contextual adaptation to news-style data while remaining computationally efficient through Low-Rank Adaptation (LoRA). Beyond performance, SentiBanglaBERT integrates SHAP-based interpretability, offering linguistic insights into how Bengali morphological cues, such as negation suffixes and aspectual markers, influence sentiment predictions. Experiments demonstrate stable performance comparable to strong baselines while providing greater transparency and interpretive depth. This framework highlights the potential of domain-adaptive continual learning as a foundation for interpretable, resource-efficient NLP in morphologically rich, underrepresented languages.

情感分类低资源语言可解释AI持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。