首个多领域约鲁巴语命名实体识别数据集,助力非洲语言NLP研究。
YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset
- 构建覆盖圣经、博客等5个领域的约鲁巴语实体标注数据集。
- 跨领域实验显示新闻与维基百科间迁移效果最好,博客电影域表现差。
- 推出约鲁巴语专用模型OyoBERT,优于通用多语言模型。
命名实体识别(NER)是自然语言处理的基础任务,但约鲁巴语研究受限于资源匮乏且领域单一。现有资源如MasakhaNER(新闻领域人工标注)和WikiAnn(维基百科自动构建)虽有价值,但覆盖范围有限。为此,我们提出YoNER,一个涵盖五个领域(圣经、博客、电影、广播、维基百科)的新型多领域约鲁巴语NER数据集,共约5,000句、100,000词,标注人物(PER)、组织(ORG)和地点(LOC)三类实体,采用CoNLL标准。由三位母语者手动标注,标注一致性超过0.70。在跨领域实验中,基于MasakhaNER 2.0评估多种Transformer编码器,同时测试少量领域内数据及跨语言设置下的表现。结果表明,非洲语系模型在约鲁巴语上优于通用多语言模型,但跨领域性能显著下降,尤其在博客和电影领域;相近正式领域(如新闻与维基)迁移更有效。此外,我们推出了新约鲁巴语专用预训练模型OyoBERT,其在本领域评测中表现优于多语言模型。数据集与OyoBERT模型已公开发布,以支持未来约鲁巴语自然语言处理研究。
原文摘要 · Abstract (English)
Named Entity Recognition (NER) is a foundational NLP task, yet research in Yorùbá has been constrained by limited and domain-specific resources. Existing resources, such as MasakhaNER (a manually annotated news-domain corpus) and WikiAnn (automatically created from Wikipedia), are valuable but restricted in domain coverage. To address this gap, we present YoNER, a new multidomain Yorùbá NER dataset that extends entity coverage beyond news and Wikipedia. The dataset comprises about 5,000 sentences and 100,000 tokens collected from five domains including Bible, Blogs, Movies, Radio broadcast and Wikipedia, and annotated with three entity types: Person (PER), Organization (ORG) and Location (LOC), following CoNLL-style guidelines. Annotation was conducted manually by three native Yorùbá speakers, with an inter-annotator agreement of over 0.70, ensuring high quality and consistency. We benchmark several transformer encoder models using cross-domain experiments with MasakhaNER 2.0, and we also assess the effect of few-shot in-domain data using YoNER and cross-lingual setups with English datasets. Our results show that African-centric models outperform general multilingual models for Yorùbá, but cross-domain performance drops substantially, particularly for blogs and movie domains. Furthermore, we observed that closely related formal domains, such as news and Wikipedia, transfer more effectively. In addition, we introduce a new Yorùbá-specific language model (OyoBERT) that outperforms multilingual models in in-domain evaluation. We publicly release the YoNER dataset and pretrained OyoBERT models to support future research on Yorùbá natural language processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。