构建首个豪萨语影评数据集,助力低资源语言情感分析
HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language
- 收集5000条豪萨语与英语混用的影评评论并三重标注
- 决策树模型达89.72%准确率,超越BERT等深度学习模型
- 证明特征工程在低资源场景下仍可实现顶尖性能
低资源语言的自然语言处理工具发展受限于标注数据匮乏。本文提出HausaMovieReview,一个包含5000条豪萨语及英豪混合语的YouTube影评评论基准数据集。数据由三位独立标注者标注,标注一致性良好,Fleiss' Kappa得分为0.85。我们对比了逻辑回归、决策树、K近邻等传统模型与微调后的BERT和RoBERTa模型。结果表明,决策树模型在准确率(89.72%)和F1分数(89.60%)上显著优于深度学习模型。研究揭示,在低资源环境下,精心设计的特征工程足以使经典模型达到先进水平,为后续研究提供了可靠基线。
原文摘要 · Abstract (English)
The development of Natural Language Processing (NLP) tools for low-resource languages is critically hindered by the scarcity of annotated datasets. This paper addresses this fundamental challenge by introducing HausaMovieReview, a novel benchmark dataset comprising 5,000 YouTube comments in Hausa and code-switched English. The dataset was meticulously annotated by three independent annotators, demonstrating a robust agreement with a Fleiss' Kappa score of 0.85 between annotators. We used this dataset to conduct a comparative analysis of classical models (Logistic Regression, Decision Tree, K-Nearest Neighbors) and fine-tuned transformer models (BERT and RoBERTa). Our results reveal a key finding: the Decision Tree classifier, with an accuracy and F1-score 89.72% and 89.60% respectively, significantly outperformed the deep learning models. Our findings also provide a robust baseline, demonstrating that effective feature engineering can enable classical models to achieve state-of-the-art performance in low-resource contexts, thereby laying a solid foundation for future research. Keywords: Hausa, Kannywood, Low-Resource Languages, NLP, Sentiment Analysis
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。