arXiv:2412.09587cs.CL2024-12EMNLP被引 3

构建50+语言的标准化命名实体识别数据集,助力多语言研究

OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages

  • 统一52种语言的标注格式与实体类型名,实现跨数据集可比性
  • 涵盖36个语料库,测试多语言模型在不同语言上表现差异
  • 适合多语言NLP、命名实体识别及大模型性能评估研究者使用

我们提出 OpenNER 1.0,一个标准化的开放获取命名实体识别(NER)数据集集合。OpenNER 包含36个跨52种语言的人工标注语料库,采用不同的实体本体结构。我们修复了标注格式问题,将原始数据统一为一致的表示形式,并对实体类型名称进行标准化,使不同语料库间保持一致。该集合支持多语言和多本体场景下的研究。我们使用三种预训练多语言模型和两种大语言模型(LLM)提供基线结果,用于比较近期模型在跨语言任务中的表现。结果表明,无单一模型在所有语言中均最优,且大模型在NER任务上的性能仍有较大提升空间。OpenNER 已开源,地址为 https://github.com/bltlab/open-ner。

原文摘要 · Abstract (English)

We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and multi-ontology NER. We provide baseline results using three pretrained multilingual language models and two large language models to compare the performance of recent models and facilitate future research in NER. We find that no single model is best in all languages and that significant work remains to obtain high performance from LLMs on the NER task. OpenNER is released at https://github.com/bltlab/open-ner.

命名实体识别多语言数据集开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。