arXiv:2508.03767cs.DBcs.IR2025-08

解决企业级超大规模数据去重与链接问题,处理1570万条记录仍高效准确。

A Robust and Efficient Pipeline for Enterprise-Level Large-Scale Entity Resolution

  • 基于AI设计可扩展的实体消歧流水线,应对高并发数据挑战。
  • 在1570万条记录上实现更高F1分数,优于传统工具如Dedupe和Splink。
  • 适合需要高可靠性的企业数据治理场景,保障数据一致性。

实体消歧(ER)在数据管理中仍是重大挑战,尤其在处理大规模数据集时。本文提出MERAI(利用AI进行大规模实体消歧),一种专为大型企业级数据集设计的稳健高效流水线,用于解决高容量数据中的记录重复与链接问题。其鲁棒性与准确性已在多个大规模消歧与链接项目中验证。为评估性能,将MERAI与两个知名实体消歧库Dedupe和Splink进行对比:Dedupe因内存限制无法扩展至超过200万条记录,而MERAI成功处理高达1570万条记录,并在所有实验中均获得准确结果。实验数据表明,MERAI在匹配准确率方面优于两个基线系统,在去重与记录链接任务中均保持更高的F1得分。MERAI为大规模企业级实体消歧提供了可扩展、可靠的解决方案,确保真实应用场景中的数据完整性与一致性。

原文摘要 · Abstract (English)

Entity resolution (ER) remains a significant challenge in data management, especially when dealing with large datasets. This paper introduces MERAI (Massive Entity Resolution using AI), a robust and efficient pipeline designed to address record deduplication and linkage issues in high-volume datasets at an enterprise level. The pipeline's resilience and accuracy have been validated through various large-scale record deduplication and linkage projects. To evaluate MERAI's performance, we compared it with two well-known entity resolution libraries, Dedupe and Splink. While Dedupe failed to scale beyond 2 million records due to memory constraints, MERAI successfully processed datasets of up to 15.7 million records and produced accurate results across all experiments. Experimental data demonstrates that MERAI outperforms both baseline systems in terms of matching accuracy, with consistently higher F1 scores in both deduplication and record linkage tasks. MERAI offers a scalable and reliable solution for enterprise-level large-scale entity resolution, ensuring data integrity and consistency in real-world applications.

实体消歧大数据处理数据治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。