arXiv:2411.04316cs.CLcs.AI2024-11被引 4

构建多语言情感词典,提升非洲低资源语言翻译与情感分析性能

A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI

  • 基于大模型与可解释AI构建跨语言情感词典,融合本地语言特征
  • BERT模型实现99%准确率,显著优于传统机器学习方法
  • 适合关注非洲低资源语言AI应用的研究者与开发者

南非和刚果民主共和国(DRC)语言环境复杂,包括祖鲁语、塞佩迪语、阿非利卡语、法语、英语及茨希卢巴语(Ciluba),由于缺乏精准标注数据,给基于AI的翻译与情感分析带来挑战。本研究开发了针对法语和茨希卢巴语的多语言情感词典,并扩展至英语、阿非利卡语、塞佩迪语和祖鲁语。词典通过整合语言特定情感得分,增强情感分类的文化相关性。构建了综合性测试语料库,训练随机森林、支持向量机(SVM)、决策树及高斯朴素贝叶斯(GNB)等模型进行低资源语言(LRLs)情感预测,其中随机森林表现最佳。此外,采用双向编码器表示(BERT)这一大型语言模型进行上下文感知情感预测,准确率达99%,精确率为98%,优于其他模型。利用可解释AI(XAI)对BERT预测结果进行解析,提升了模型透明度并增强了可信度。研究结果表明,该词典与模型显著提升了南非和刚果民主共和国低资源语言的翻译与情感分析能力,为未来支持未充分代表语言的AI模型奠定基础,适用于教育、治理与商业等多语言场景。

原文摘要 · Abstract (English)

South Africa and the Democratic Republic of Congo (DRC) present a complex linguistic landscape with languages such as Zulu, Sepedi, Afrikaans, French, English, and Tshiluba (Ciluba), which creates unique challenges for AI-driven translation and sentiment analysis systems due to a lack of accurately labeled data. This study seeks to address these challenges by developing a multilingual lexicon designed for French and Tshiluba, now expanded to include translations in English, Afrikaans, Sepedi, and Zulu. The lexicon enhances cultural relevance in sentiment classification by integrating language-specific sentiment scores. A comprehensive testing corpus is created to support translation and sentiment analysis tasks, with machine learning models such as Random Forest, Support Vector Machine (SVM), Decision Trees, and Gaussian Naive Bayes (GNB) trained to predict sentiment across low resource languages (LRLs). Among them, the Random Forest model performed particularly well, capturing sentiment polarity and handling language-specific nuances effectively. Furthermore, Bidirectional Encoder Representations from Transformers (BERT), a Large Language Model (LLM), is applied to predict context-based sentiment with high accuracy, achieving 99% accuracy and 98% precision, outperforming other models. The BERT predictions were clarified using Explainable AI (XAI), improving transparency and fostering confidence in sentiment classification. Overall, findings demonstrate that the proposed lexicon and machine learning models significantly enhance translation and sentiment analysis for LRLs in South Africa and the DRC, laying a foundation for future AI models that support underrepresented languages, with applications across education, governance, and business in multilingual contexts.

情感分析低资源语言大模型可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。