arXiv:2501.11023cs.CL2025-01被引 7

用LAFT微调AfriBERTa提升豪萨语情感分析效果,推动低资源语言NLP发展。

Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in Hausa Language Using AfriBERTa

  • 通过构建无标签语料库并应用语言自适应微调,增强模型对豪萨语的适配性。
  • 在NaijaSenti数据集上,微调后模型性能优于未针对豪萨语训练的基线模型。
  • 研究强调多样数据源对非洲低资源语言NLP的重要性,适合关注公平AI与本土化NLP的研究者。

情感分析(SA)在自然语言处理中至关重要,用于识别文本中的情感倾向。尽管主流语言取得了显著进展,但像豪萨语这样的低资源语言仍面临数字资源匮乏的挑战。本研究探究语言自适应微调(LAFT)对豪萨语情感分析的效果。首先,我们构建了一个多样化且未标注的语料库以扩展模型的语言能力,随后使用LAFT将AfriBERTa模型专门适配豪萨语的语义特征。接着,在标注的NaijaSenti数据集上进行微调以评估性能。结果表明,LAFT带来适度提升,可能因采用正式豪萨语文本而非社交媒体非正式数据所致。然而,预训练的AfriBERTa模型显著优于未针对豪萨语训练的模型,凸显了预训练模型在低资源场景中的关键作用。本研究强调需使用多样数据源来推进非洲低资源语言的NLP应用。代码与数据集已公开,促进可复现性与后续研究。

原文摘要 · Abstract (English)

Sentiment analysis (SA) plays a vital role in Natural Language Processing (NLP) by ~identifying sentiments expressed in text. Although significant advances have been made in SA for widely spoken languages, low-resource languages such as Hausa face unique challenges, primarily due to a lack of digital resources. This study investigates the effectiveness of Language-Adaptive Fine-Tuning (LAFT) to improve SA performance in Hausa. We first curate a diverse, unlabeled corpus to expand the model's linguistic capabilities, followed by applying LAFT to adapt AfriBERTa specifically to the nuances of the Hausa language. The adapted model is then fine-tuned on the labeled NaijaSenti sentiment dataset to evaluate its performance. Our findings demonstrate that LAFT gives modest improvements, which may be attributed to the use of formal Hausa text rather than informal social media data. Nevertheless, the pre-trained AfriBERTa model significantly outperformed models not specifically trained on Hausa, highlighting the importance of using pre-trained models in low-resource contexts. This research emphasizes the necessity for diverse data sources to advance NLP applications for low-resource African languages. We published the code and the dataset to encourage further research and facilitate reproducibility in low-resource NLP here: https://github.com/Sani-Abdullahi-Sani/Natural-Language-Processing/blob/main/Sentiment%20Analysis%20for%20Low%20Resource%20African%20Languages

情感分析低资源语言AfriBERTa豪萨语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。