arXiv:2602.23941cs.CLcs.DL2026-02中稿 · LREC 2026

构建历史地理坐标标注集,提升古籍文本坐标的自动提取能力。

EDDA-Coordinata: An Annotated Dataset of Historical Geographic Coordinates

  • 基于启蒙时代《百科全书》文本,人工标注4798条含坐标的条目。
  • 训练的模型在法语古籍上达到61%准确率,英文近代百科达77%。
  • 数据集与方法可跨语言、跨时代用于历史文献数字化研究。

本文介绍了一个从狄德罗与达朗贝尔十八世纪《百科全书》中提取并增强的地理坐标数据集。从ARTFL和ENCCRE两个数字化版本共7.4万篇文章中,筛选出15,278条地理条目,人工识别出4,798条包含坐标的数据,另有10,480条为非数值描述。基于此金标准标注,我们训练了基于Transformer的模型以实现坐标检索与归一化。所提流程结合分类器识别含坐标条目及第二阶段坐标提取模型,测试了编码器-解码器与解码器架构。交叉验证得86%精确匹配(EM)分数。在法语十八世纪《特雷沃辞典》上,微调模型达61% EM;在英语十九世纪第七版《大英百科全书》上达77%。结果表明该数据集具强训练价值,且两步法具备跨语言、跨领域泛化能力。

原文摘要 · Abstract (English)

This paper introduces a dataset of enriched geographic coordinates retrieved from Diderot and d'Alembert's eighteenth-century Encyclopedie. Automatically recovering geographic coordinates from historical texts is a complex task, as they are expressed in a variety of ways and with varying levels of precision. To improve retrieval of coordinates from similar digitized early modern texts, we have created a gold standard dataset, trained models, published the resulting inferred and normalized coordinate data, and experimented applying these models to new texts. From 74,000 total articles in each of the digitized versions of the Encyclopedie from ARTFL and ENCCRE, we examined 15,278 geographical entries, manually identifying 4,798 containing coordinates, and 10,480 with descriptive but non-numerical references. Leveraging our gold standard annotations, we trained transformer-based models to retrieve and normalize coordinates. The pipeline presented here combines a classifier to identify coordinate-bearing entries and a second model for retrieval, tested across encoder-decoder and decoder architectures. Cross-validation yielded an 86% EM score. On an out-of-domain eighteenth-century Trevoux dictionary (also in French), our fine-tuned model had a 61% EM score, while for the nineteenth-century, 7th edition of the Encyclopaedia Britannica in English, the EM was 77%. These findings highlight the gold standard dataset's usefulness as training data, and our two-step method's cross-lingual, cross-domain generalizability.

历史地理文本挖掘标注数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。