首个针对孟加拉语的情感检测基准,提出高效模型与数据集。
EmoBang: Detecting Emotion From Bengali Texts
- 构建8类情绪标注的孟加拉语数据集,填补研究空白。
- 提出CRNN与BERT集成模型,准确率分别达92.86%和93.69%。
- 首次系统评估大模型在低资源语言上的零样本能力,适合多语言研究者。
文本情感检测旨在根据语言线索识别个体的情绪状态(积极、消极或中性)。尽管英语等高资源语言已取得显著进展,孟加拉语作为全球第四大语言却仍被严重忽视,缺乏大规模标准化数据集,属于低资源语言。现有研究多依赖传统机器学习与特征工程,性能有限。本文提出一个涵盖八类情绪的新孟加拉语情感数据集,并设计两种自动情感检测模型:(i) 混合卷积循环神经网络模型(EmoBangHybrid),(ii) AdaBoost-BERT集成模型(EmoBangEnsemble)。此外,我们评估了六种基线模型及五种特征工程方法,并测试零样本与少量样本大语言模型在该数据集上的表现。据我们所知,这是首个全面的孟加拉语情感检测基准。实验结果表明,EmoBangH与EmoBangE的准确率分别为92.86%和93.69%,优于现有方法,为未来研究建立了强基线。
原文摘要 · Abstract (English)
Emotion detection from text seeks to identify an individual's emotional or mental state - positive, negative, or neutral - based on linguistic cues. While significant progress has been made for English and other high-resource languages, Bengali remains underexplored despite being the world's fourth most spoken language. The lack of large, standardized datasets classifies Bengali as a low-resource language for emotion detection. Existing studies mainly employ classical machine learning models with traditional feature engineering, yielding limited performance. In this paper, we introduce a new Bengali emotion dataset annotated across eight emotion categories and propose two models for automatic emotion detection: (i) a hybrid Convolutional Recurrent Neural Network (CRNN) model (EmoBangHybrid) and (ii) an AdaBoost-Bidirectional Encoder Representations from Transformers (BERT) ensemble model (EmoBangEnsemble). Additionally, we evaluate six baseline models with five feature engineering techniques and assess zero-shot and few-shot large language models (LLMs) on the dataset. To the best of our knowledge, this is the first comprehensive benchmark for Bengali emotion detection. Experimental results show that EmoBangH and EmoBangE achieve accuracies of 92.86% and 93.69%, respectively, outperforming existing methods and establishing strong baselines for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。