arXiv:2411.00030cs.CLcs.AI2024-11被引 1

修复法语命名实体识别数据集的标注质量,构建高质量标准语料

WikiNER-fr-gold: A Gold-Standard NER Corpus

  • 基于原始法语语料,人工复核并修正标注错误
  • 保留20%原始数据(26,818句,70万词元)作为新语料
  • 适合需要高精度法语NER训练的研究者使用

本文针对多语言命名实体识别语料WikiNER的质量问题,提出其法语部分的修订版本WikiNER-fr-gold。原语料采用半监督方式生成,未进行后期人工校验,属于银标准数据集。本工作从梳理各类实体类型出发制定标注规范,对随机抽取的原法语子集(26,818句,700k词元)进行人工复核与修正。通过分析原始语料中的错误与不一致性,构建出更可靠的金标准语料,并探讨未来改进方向。

原文摘要 · Abstract (English)

We address in this article the the quality of the WikiNER corpus, a multilingual Named Entity Recognition corpus, and provide a consolidated version of it. The annotation of WikiNER was produced in a semi-supervised manner i.e. no manual verification has been carried out a posteriori. Such corpus is called silver-standard. In this paper we propose WikiNER-fr-gold which is a revised version of the French proportion of WikiNER. Our corpus consists of randomly sampled 20% of the original French sub-corpus (26,818 sentences with 700k tokens). We start by summarizing the entity types included in each category in order to define an annotation guideline, and then we proceed to revise the corpus. Finally we present an analysis of errors and inconsistency observed in the WikiNER-fr corpus, and we discuss potential future work directions.

命名实体识别数据集法语金标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。