用用户生成的肯尼亚混合语推文,提升低资源语言的情感与情绪识别效果。
RideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched Dataset
- 基于真实推文构建肯尼亚多语混合数据集,采用监督与半监督方法评估模型。
- XLM-R在情感分析中表现最佳,准确率达69.2%,F1达66.1%。
- 发现模型普遍倾向预测中性,且非洲本地模型对共情情绪敏感。
社交媒体已成为人们表达观点和分享经历的重要开放平台。然而,由于内容稀缺、质量差以及语言使用差异大(如俚语和语言混杂),利用低资源语言的推文数据极具挑战性。本文分析肯尼亚代码切换数据,评估四种先进Transformer预训练模型在情感与情绪分类中的表现,采用监督与半监督方法。详细阐述了数据收集与标注方法,并指出数据清洗阶段面临的困难。结果表明,XLM-R在情感分析中表现最优:其监督模型准确率达到69.2%,F1得分为66.1%;半监督模型准确率为67.2%,F1为64.1%。情绪分析中,DistilBERT监督模型准确率最高(59.8%),F1为31%;mBERT半监督模型准确率59%,F1为26.5%。AfriBERTa系列模型表现最差。所有模型均倾向于预测中性情绪,其中AfriBERT表现出最强偏差,并对‘共情’情绪具有独特敏感性。
原文摘要 · Abstract (English)
Social media has become a crucial open-access platform for individuals to express opinions and share experiences. However, leveraging low-resource language data from Twitter is challenging due to scarce, poor-quality content and the major variations in language use, such as slang and code-switching. Identifying tweets in these languages can be difficult as Twitter primarily supports high-resource languages. We analyze Kenyan code-switched data and evaluate four state-of-the-art (SOTA) transformer-based pretrained models for sentiment and emotion classification, using supervised and semi-supervised methods. We detail the methodology behind data collection and annotation, and the challenges encountered during the data curation phase. Our results show that XLM-R outperforms other models; for sentiment analysis, XLM-R supervised model achieves the highest accuracy (69.2\%) and F1 score (66.1\%), XLM-R semi-supervised (67.2\% accuracy, 64.1\% F1 score). In emotion analysis, DistilBERT supervised leads in accuracy (59.8\%) and F1 score (31\%), mBERT semi-supervised (accuracy (59\% and F1 score 26.5\%). AfriBERTa models show the lowest accuracy and F1 scores. All models tend to predict neutral sentiment, with Afri-BERT showing the highest bias and unique sensitivity to empathy emotion. https://github.com/NEtori21/Ride_hailing
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。