arXiv:2512.01557cs.CL2025-12

非洲语言在数字空间中数据稀缺,新闻文本比社交平台更适合作为训练AI的可靠语料。

Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

  • 用新闻媒体数据替代社交平台,获取更纯净的非洲语言语料
  • 新闻数据上语言检测模型准确率接近100%,社交平台则严重下降
  • 提示需开发能处理混用语言的模型,助力非洲语言AI发展

本研究考察非洲语言在数字空间中的表现及其对当前语言识别工具带来的挑战。我们评估了尤鲁巴语、基尼亚鲁旺达语和阿姆哈拉语的表现。尽管这些语言有数百万使用者,其在对话类平台上的在线使用量稀少,且高度受英语影响,无法代表母语者之间的纯语言交流。这种真实对话数据的匮乏,导致训练语言模型时可用数据稀缺。为此,我们从Reddit子版块和本地新闻来源收集了每种语言的数据。分析显示两类来源差异显著:Reddit数据极少且频繁出现语言混杂;而本地新闻媒体提供了丰富、纯净的单语数据,并促使新闻机构社交媒体页面上的用户更多使用本地语言互动。语言检测模型(包括宏观分类器GlotLID、专用模型AfroLID及通用大模型Llama 3.3 70B)在干净新闻数据上表现近乎完美,但在混用语言的Reddit内容中表现不佳。研究结论认为,专业编排的新闻内容是训练上下文丰富型非洲语言AI模型更可靠、有效的数据来源,同时强调未来需开发能处理纯净与混用文本的模型以提升非洲语言识别准确率。

原文摘要 · Abstract (English)

This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evaluate their performance on Yoruba, Kinyarwanda, and Amharic. While these languages are spoken by millions, their online usage on conversational platforms is often sparse, heavily influenced by English, and not representative of the authentic, monolingual conversations prevalent among native speakers. This lack of readily available authentic data online creates a challenge of scarcity of conversational data for training language models. To investigate this, data was collected from subreddits and local news sources for each language. The analysis showed a stark contrast between the two sources. Reddit data was minimal and characterized by heavy code-switching. Conversely, local news media offered a robust source of clean, monolingual language data, which also prompted more user engagement in the local language on the news publishers' social media pages. Language detection models, including a macro-classifier (GlotLID), the specialized AfroLID, and a general-purpose LLM (Llama 3.3 70B), performed with near-perfect accuracy on the clean news data but struggled with the code-switched Reddit posts. The study concludes that professionally curated news content is a more reliable and effective source for training context-rich AI models for African languages than data from conversational platforms. It also highlights the need for future models that can process clean and code-switched text to improve the detection accuracy for African languages.

非洲语言语言检测数据稀缺新闻语料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。