arXiv:2601.21138cs.CL2026-01被引 1

无需标注数据,用预训练模型实现高精度记录链接

EnsembleLink: Accurate Record Linkage Without Training Data

  • 利用预训练语言模型捕捉语义关系进行链接
  • 在多类数据上达到或超越需大量标注的方法
  • 本地运行、快速高效,适合无标签场景

记录链接是跨数据集匹配同一实体的关键步骤,对实证社会科学至关重要,但方法论仍不成熟。研究者常将其作为预处理,采用随意规则,却未量化链接错误对后续分析的影响。现有方法要么准确率低,要么需大量标注数据。本文提出EnsembleLink,一种无需任何训练标签即可实现高准确率的链接方法。该方法利用预训练语言模型从大规模文本中学习语义关系(如“南奥泽公园”是“纽约市”的一个街区,或“工人斗争”指托洛茨基主义的“工人斗争”党)。在涵盖城市名、人名、组织、多语种政党及文献记录的基准测试中,EnsembleLink的表现达到或超过需大量标注的方法。该方法可在本地使用开源模型运行,无需外部API调用,典型任务可在数分钟内完成。

原文摘要 · Abstract (English)

Record linkage, the process of matching records that refer to the same entity across datasets, is essential to empirical social science but remains methodologically underdeveloped. Researchers treat it as a preprocessing step, applying ad hoc rules without quantifying the uncertainty that linkage errors introduce into downstream analyses. Existing methods either achieve low accuracy or require substantial labeled training data. I present EnsembleLink, a method that achieves high accuracy without any training labels. EnsembleLink leverages pre-trained language models that have learned semantic relationships (e.g., that "South Ozone Park" is a neighborhood in "New York City" or that "Lutte ouvriere" refers to the Trotskyist "Workers' Struggle" party) from large text corpora. On benchmarks spanning city names, person names, organizations, multilingual political parties, and bibliographic records, EnsembleLink matches or exceeds methods requiring extensive labeling. The method runs locally on open-source models, requiring no external API calls, and completes typical linkage tasks in minutes.

记录链接预训练模型无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。