arXiv:2601.13346cs.CL2026-01被引 1

构建非洲语言识别框架,覆盖640种语言并提升细粒度辨识能力

AfroScope: A Framework for Studying the Linguistic Landscape of Africa

  • 提出分层分类方法,用专用嵌入模型解决近似语言混淆问题
  • 在易混淆语言子集上,宏平均F1提升1.57点,达当前最优
  • 适合关注非洲语言技术、多语种NLP的研究者与开发者

语言识别(LID)是影响下游自然语言处理应用可靠性的基础预处理步骤。尽管近期研究扩展了非洲语言的LID能力,现有系统在语言覆盖和近似语言/方言的细粒度区分上仍存在局限。本文提出AfroScope统一框架,包含覆盖640种语言的AfroScope-Data数据集,以及具备广泛非洲语言覆盖的AfroScope-Models模型套件。为解决近似语言间的持续混淆问题,提出一种分层分类方法,利用专用于目标消歧的AfroScope-Mirror嵌入模型,在易混淆语言子集上相较最佳基线模型,宏平均F1提升1.57点。进一步分析跨语言迁移与领域效应,揭示语言家族结构、文字兼容性及领域覆盖对LID性能的影响。将非洲语言识别定位为大规模数字化文本中测量非洲语言景观的关键技术,并已公开发布AfroScope-Data与AfroScope-Models。

原文摘要 · Abstract (English)

Language Identification (LID), the task of determining the language of a given text, is a fundamental preprocessing step that shapes the reliability of downstream NLP applications. While recent work has expanded African LID, existing systems remain limited in both language coverage and fine-grained discrimination among closely related languages and varieties. We introduce AfroScope, a unified framework for African LID that includes AfroScope-Data, a dataset covering 640 languages, and AfroScope-Models, a suite of strong LID models with broad African language coverage. To address persistent confusions among closely related languages, we propose a hierarchical classification approach that leverages AfroScope-Mirror, a specialized embedding model for targeted disambiguation, improving macro-F1 by 1.57 points on the confusable subset compared to our best base model. We further analyze cross-lingual transfer and domain effects, showing how language-family structure, script compatibility, and domain coverage shape LID performance. We position African LID as an enabling technology for large-scale measurement of Africa's linguistic landscape in digital text, and release AfroScope-Data and AfroScope-Models online.

语言识别非洲语言多语种NLP嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。