arXiv:2603.11051cs.IRcs.AI2026-03

首个大规模制裁实体匹配基准,助力合规数据精准识别。

OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

  • 构建涵盖百万实体的多语言、跨系统实体匹配数据集
  • GPT-4o达99.0% F1,开源模型达98.2%性能接近实用上限
  • 揭示规则与LLM互补缺陷,推动匹配流程优化

我们发布OpenSanctions Pairs,首个针对制裁与公开情报(OSINT)数据的大规模公开实体匹配基准。数据集包含超过100万实体的755,540对专家标注样本,源自45个司法管辖区的293个数据源。其真实世界多样性覆盖多种语言和书写系统(如拉丁、西里尔、阿拉伯文),结构不一致且随时间变化,显著高于以往实体匹配基准。作为基线,我们评估了生产级规则匹配器(nomenklatura RegressionV1)及开源与闭源LLMs在零样本和少样本设置下的表现,并结合或不结合MIPROv2提示优化以控制提示敏感性。规则基线达到91.3% F1;GPT-4o取得最佳成绩99.0% F1,本地部署的开源模型DeepSeek-R1-Distill-Qwen-14B达到98.2% F1。规则方法过匹配,而LLMs在跨脚本音译上表现不佳,二者缺陷互补。结果表明,配对匹配性能已趋近实用上限,需将关注点转向阻断、聚类和不确定性感知审查等管道组件。

原文摘要 · Abstract (English)

We release OpenSanctions Pairs, the first large-scale public benchmark for entity matching on sanctions and OSINT data. The dataset includes 755,540 expert-labeled pairs over 1 million entities, aggregated from 293 source datasets across 45 jurisdictions. It captures real-world diversity in compliance data, spanning multiple languages and writing systems (e.g., Latin, Cyrillic, Arabic), inconsistent structure, and time-varying provenance, and is substantially more heterogeneous than prior entity matching benchmarks. As baselines, we evaluate the production rule-based matcher (nomenklatura RegressionV1) alongside open- and closed-source LLMs in both zero- and few-shot settings, each tested with and without MIPROv2 prompt optimization to control for prompt sensitivity. The rule-based baseline reaches 91.3\% F1; GPT-4o achieves the best result at 99.0\% F1, and a locally deployable open-source model (DeepSeek-R1-Distill-Qwen-14B) achieves 98.2\% F1. The rule-based baseline and LLMs fail in complementary ways: rules over-match, while LLMs struggle with cross-script transliteration. These results suggest that pairwise matching performance is approaching a practical ceiling and shift attention toward pipeline components such as blocking, clustering, and uncertainty-aware review.

实体匹配大模型合规数据多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。