arXiv:2506.13792cs.AIcs.CL2025-06

构建了覆盖220年的冰岛人口普查数据集,用于长期身份匹配研究。

ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution

  • 整合220年冰岛普查数据,含多层级地理与亲属关系信息
  • 涵盖98.4万条记录,22.7万条人工标注身份标识
  • 适用于历史人物识别、跨时代实体匹配等场景

我们提出ICE-ID,一个包含16个冰岛人口普查周期(1703–1920年)共984,028条记录的基准数据集,其中226,864个身份标识由专家手工标注。该数据集融合了层级地理结构(农场→教区→地区→县)、父名命名惯例、稀疏亲属关系(配偶、父亲、母亲)以及跨几十年的时间漂移,这些挑战在标准产品匹配或引文数据集中均未体现。本文对时间覆盖范围、缺失性、标识歧义、候选生成效率和聚类分布进行了基于实物分析,并将ICE-ID与经典实体解析基准(Abt-Buy、Amazon-Google、DBLP-ACM、DBLP-Scholar、Walmart-Amazon、iTunes-Amazon、Beer、Fodors-Zagats)进行对比。同时定义了符合部署实际的时间外样本(OOD)评估协议,公开发布数据集、划分方案、再生脚本、分析成果及交互式探索仪表盘。基线模型比较与端到端实体解析结果详见配套方法论文。

原文摘要 · Abstract (English)

We introduce \textbf{ICE-ID}, a benchmark dataset comprising 984,028 records from 16 Icelandic census waves spanning 220 years (1703--1920), with 226,864 expert-curated person identifiers. ICE-ID combines hierarchical geography (farm$\to$parish$\to$district$\to$county), patronymic naming conventions, sparse kinship links (partner, father, mother), and multi-decadal temporal drift -- challenges not captured by standard product-matching or citation datasets. This paper presents an artifact-backed analysis of temporal coverage, missingness, identifier ambiguity, candidate-generation efficiency, and cluster distributions, and situates ICE-ID against classical ER benchmarks (Abt--Buy, Amazon--Google, DBLP--ACM, DBLP--Scholar, Walmart--Amazon, iTunes--Amazon, Beer, Fodors--Zagats). We also define a deployment-faithful temporal OOD protocol and release the dataset, splits, regeneration scripts, analysis artifacts, and a dashboard for interactive exploration. Baseline model comparisons and end-to-end ER results are reported in the companion methods paper.

历史数据实体解析身份匹配数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。