构建16种阿拉伯语方言跨10个领域的标注语料库,用于评估命名实体识别模型性能。
Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition
- 构建160个跨方言与领域的标注语料,覆盖77.7万词元
- 跨领域/方言测试中模型性能最高下降38%
- 开源共享,适用于方言适应与迁移学习研究
我们提出Konooz,一个涵盖16种阿拉伯语方言、10个领域的多维度语料库,共形成160个独立语料。该语料包含约77.7万词元,经人工标注,采用Wojood标准进行21类实体的嵌套与平铺标注。尽管Konooz可用于多种NLP任务如领域自适应与迁移学习,本文主要聚焦于基准测试现有阿拉伯语命名实体识别(NER)模型,尤其是跨领域与跨方言表现。使用Konooz对四种阿拉伯语NER模型的评测显示,相较于分布内数据,性能最高下降38%。我们还通过最大均值差异(MMD)度量分析了领域与方言间的差异,并揭示资源稀缺性对性能的影响。某些模型在特定方言或领域表现更优的原因也得以阐释。Konooz已开源,可通过https://sina.birzeit.edu/wojood/#download 获取。
原文摘要 · Abstract (English)
We introduce Konooz, a novel multi-dimensional corpus covering 16 Arabic dialects across 10 domains, resulting in 160 distinct corpora. The corpus comprises about 777k tokens, carefully collected and manually annotated with 21 entity types using both nested and flat annotation schemes - using the Wojood guidelines. While Konooz is useful for various NLP tasks like domain adaptation and transfer learning, this paper primarily focuses on benchmarking existing Arabic Named Entity Recognition (NER) models, especially cross-domain and cross-dialect model performance. Our benchmarking of four Arabic NER models using Konooz reveals a significant drop in performance of up to 38% when compared to the in-distribution data. Furthermore, we present an in-depth analysis of domain and dialect divergence and the impact of resource scarcity. We also measured the overlap between domains and dialects using the Maximum Mean Discrepancy (MMD) metric, and illustrated why certain NER models perform better on specific dialects and domains. Konooz is open-source and publicly available at https://sina.birzeit.edu/wojood/#download
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。