首个专为索马里语设计的Transformer模型,提升假新闻与有害信息检测能力。
Detection of Somali-written Fake News and Toxic Messages on the Social Media Using Transformer-based Language Models
- 基于社交媒体数据构建两个索马里语标注数据集,训练出首个单语索马里语模型SomBERTa。
- 在假新闻和有毒内容分类任务中,平均准确率达87.99%,优于多语言模型。
- 为低资源语言提供可复用的NLP框架,推动数字包容与语言多样性。
社交媒体使人人可发布内容,公众对平台信息依赖加深,导致虚假新闻、有害内容等问题加剧。尽管人工审核有一定作用,但AI模型更具可持续性与可扩展性。然而,索马里语等低资源语言面临标注数据匮乏、专用语言模型缺失等问题。本文提出部分持续研究工作,针对索马里语构建两个面向假新闻与毒性内容分类的人工标注数据集,并开发首个单语索马里语Transformer模型SomBERTa。该模型在毒性内容、假新闻及新闻主题分类数据集上进行微调与评估。与AfriBERTa、AfroXLMR等多语言模型对比显示,SomBERTa在假新闻与毒性内容分类任务中均表现更优,所有任务平均准确率达87.99%。本研究为索马里语NLP提供了基础模型与可复现框架,促进数字与AI包容性及语言多样性。
原文摘要 · Abstract (English)
The fact that everyone with a social media account can create and share content, and the increasing public reliance on social media platforms as a news and information source bring about significant challenges such as misinformation, fake news, harmful content, etc. Although human content moderation may be useful to an extent and used by these platforms to flag posted materials, the use of AI models provides a more sustainable, scalable, and effective way to mitigate these harmful contents. However, low-resourced languages such as the Somali language face limitations in AI automation, including scarce annotated training datasets and lack of language models tailored to their unique linguistic characteristics. This paper presents part of our ongoing research work to bridge some of these gaps for the Somali language. In particular, we created two human-annotated social-media-sourced Somali datasets for two downstream applications, fake news \& toxicity classification, and developed a transformer-based monolingual Somali language model (named SomBERTa) -- the first of its kind to the best of our knowledge. SomBERTa is then fine-tuned and evaluated on toxic content, fake news and news topic classification datasets. Comparative evaluation analysis of the proposed model against related multilingual models (e.g., AfriBERTa, AfroXLMR, etc) demonstrated that SomBERTa consistently outperformed these comparators in both fake news and toxic content classification tasks while achieving the best average accuracy (87.99%) across all tasks. This research contributes to Somali NLP by offering a foundational language model and a replicable framework for other low-resource languages, promoting digital and AI inclusivity and linguistic diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。