首个面向历史意大利语的实体识别与链接公开数据集,覆盖两百年文献。
ENEIDE: A High Quality Silver Standard Dataset for Named Entity Recognition and Linking in Historical Italian
- 从两个学术数字版本中半自动提取实体标注,含质量控制流程。
- 包含超8000个实体标注,覆盖人物、地点、组织等类型并链接维基数据。
- 适合研究历史文本中的时序实体消歧与跨领域模型评估。
本文提出ENEIDE(从意大利数字版中提取命名实体),首个面向历史意大利语的命名实体识别与链接(NERL)银标准数据集。该语料库包含2,111篇文档,来自两个学术数字版本:哲学家贾科莫·莱奥帕尔迪的《数字齐巴尔多内》(1798–1837)和政治家阿尔多·莫罗的《阿尔多·莫罗数字版》(1916–1978)。共标注超过8,000个实体,涵盖人物、地点、组织、文学作品等类型,并与维基数据标识符关联,包括无法映射至知识图谱的NIL实体。据我们所知,ENEIDE是首个多领域、公开可用的意大利历史语料库,提供训练、开发和测试划分。论文提出从人工校订数字版本中半自动提取标注的方法,包含质量控制与标注增强流程。基于先进模型的基线实验表明,该数据集对NERL构成挑战,零样本方法与微调模型间存在显著差距。其横跨两个世纪的时间跨度,使其特别适用于时序实体消歧与跨领域评估。ENEIDE采用CC BY-NC-SA 4.0许可发布。
原文摘要 · Abstract (English)
This paper introduces ENEIDE (Extracting Named Entities from Italian Digital Editions), a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 documents with over 8,000 entity annotations semi-automatically extracted from two scholarly digital editions: Digital Zibaldone, the philosophical diary of the Italian poet Giacomo Leopardi (1798--1837), and Aldo Moro Digitale, the complete works of the Italian politician Aldo Moro (1916--1978). Annotations cover multiple entity types (person, location, organization, literary work) linked to Wikidata identifiers, including NIL entities that cannot be mapped to the knowledge graph. To the best of our knowledge, ENEIDE represents the first multi-domain, publicly available NERL dataset for historical Italian with training, development, and test splits. We present a methodology for semi-automatic annotations extraction from manually curated scholarly digital editions, including quality control and annotation enhancement procedures. Baseline experiments using state-of-the-art models demonstrate the dataset's challenge for NERL and the gap between zero-shot approaches and fine-tuned models. The dataset's diachronic coverage spanning two centuries makes it particularly suitable for temporal entity disambiguation and cross-domain evaluation. ENEIDE is released under a CC BY-NC-SA 4.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。