arXiv:2512.02799cs.CL2025-12

构建多语言情感词典框架,提升南非低资源语言的情感分析性能

TriLex: A Framework for Multilingual Sentiment Analysis in Low-Resource South African Languages

  • 三阶段框架融合语料抽取、跨语言映射与检索生成,扩展低资源语言词典
  • AfroXLMR在isiXhosa和isiZulu上F1超80%,跨语言稳定性强
  • 适用于资源受限环境,适合研究非洲语言NLP的学者与开发者

低资源非洲语言在情感分析中仍被严重忽视,制约了词汇覆盖与多语言自然语言处理系统的性能。本文提出TriLex,一种三阶段检索增强框架,通过语料库提取、跨语言映射和检索增强生成(RAG)驱动的词汇精炼,系统性扩展低资源语言的情感词典。利用丰富后的词典,评估了两种主流非洲预训练语言模型(AfroXLMR 和 AfriBERTa)在多个案例中的表现。结果表明,AfroXLMR表现更优,在isiXhosa和isiZulu上F1分数超过80%,展现出强跨语言稳定性;尽管AfriBERTa未在目标语言上预训练,其F1分数仍稳定在约64%,验证了其在计算资源受限场景下的实用性。两种模型均优于传统机器学习基线,集成分析进一步提升了精度与鲁棒性。研究证明TriLex是低资源南非语言情感词典扩展与建模的有效可扩展框架。

原文摘要 · Abstract (English)

Low-resource African languages remain underrepresented in sentiment analysis, limiting both lexical coverage and the performance of multilingual Natural Language Processing (NLP) systems. This study proposes TriLex, a three-stage retrieval augmented framework that unifies corpus-based extraction, cross lingual mapping, and retrieval augmented generation (RAG) driven lexical refinement to systematically expand sentiment lexicons for low-resource languages. Using the enriched lexicon, the performance of two prominent African pretrained language models (AfroXLMR and AfriBERTa) is evaluated across multiple case studies. Results demonstrate that AfroXLMR delivers superior performance, achieving F1-scores above 80% for isiXhosa and isiZulu and exhibiting strong cross-lingual stability. Although AfriBERTa lacks pre-training on these target languages, it still achieves reliable F1-scores around 64%, validating its utility in computationally constrained settings. Both models outperform traditional machine learning baselines, and ensemble analyses further enhance precision and robustness. The findings establish TriLex as a scalable and effective framework for multilingual sentiment lexicon expansion and sentiment modeling in low-resource South African languages.

情感分析多语言低资源语言非洲语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。