构建首个大规模多语言命名实体识别基准,支持100+语言的标准化评估。
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark

- 采用统一标签体系与标注规范,实现跨语言实体标注标准化。
- 覆盖100多种语言,包含数百万条高质量标注数据。
- 适合多语言NLP研究者、模型评估者及跨语言应用开发者使用。
尽管多语言大模型有望为众多语言使用者带来大模型优势,但多数语言缺乏高质量的评估基准来验证这些假设。Universal NER项目已进入第四年,致力于构建高质量的多语言命名实体识别(NER)基准数据集。受其他核心NLP任务(如Universal Dependencies)大规模多语言工作的启发,项目采用通用标签集和详尽的标注指南,收集标准化的跨语言实体跨度标注。首个版本(UNER v1)于2024年发布,此后项目持续扩展,汇聚了众多组织者、标注员与合作者,形成活跃社区。
原文摘要 · Abstract (English)
While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these assumptions remain scarce. The Universal NER project, now entering its fourth year, is dedicated to building gold-standard multilingual Named Entity Recognition (NER) benchmark datasets. Inspired by existing massively multilingual efforts for other core NLP tasks (e.g., Universal Dependencies), the project uses a general tagset and thorough annotation guidelines to collect standardized, cross-lingual annotations of named entity spans. The first installment (UNER v1) was released in 2024, and the project has continued and expanded since then, with various organizers, annotators, and collaborators in an active community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。