首个非洲语言句法树库集合,助力自然语言处理研究
AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages

- 构建九种非洲语言的句法标注语料库,基于通用依赖框架
- 模型在九种语言上表现普遍不足,暴露语法理解鸿沟
- 适合关注非洲语言、跨语言泛化与低资源语言研究者
尽管非洲语言具有语言多样性且全球意义重大,但在自然语言处理研究和资源支持方面仍严重不足。本文提出 AfriSUD,首个涵盖九种不同非洲语言的大规模句法标注树库集合,覆盖撒哈拉以南非洲主要语系和区域。基于表面句法通用依存(SUD)框架,该数据集由母语者验证,高质量地捕捉了黏着性、声调等类型学关键特征。我们在 AfriSUD 上评估了多种模型在词性标注和依存句法分析任务上的表现,包括非变换器基线、多语言预训练编码器及大语言模型。结果揭示出显著的语法差距:各类模型在九种语言上均存在明显局限,表明现有架构可能无法充分捕捉非洲语言语法结构的多样性。
原文摘要 · Abstract (English)
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。