用分割学习实现隐私保护的记录匹配,效果接近传统方法。
Towards Split Learning-based Privacy-Preserving Record Linkage
- 引入公开参考集训练分割学习模型,避免数据泄露。
- 在真实数据上匹配准确率损失小于1.5%,接近中心化SVM。
- 适合医疗、金融等需跨机构数据匹配的隐私敏感场景。
分割学习近年来被提出以支持用户数据隐私保护的应用。然而,该技术在隐私保护记录匹配(Privacy-Preserving Record Linkage)中的应用尚未得到充分研究——该问题要求在不同数据持有者数据库中识别同一实体,同时不泄露额外信息。本文通过引入基于公开参考集(Reference Sets)的新训练方法,探索了分割学习在隐私保护记录匹配中的潜力。实验表明,该方法在真实数据集上的匹配性能与传统的集中式支持向量机(SVM)方法相比,仅产生低于1.5%的准确率下降,且无需共享原始数据,有效保障了隐私。
原文摘要 · Abstract (English)
Split Learning has been recently introduced to facilitate applications where user data privacy is a requirement. However, it has not been thoroughly studied in the context of Privacy-Preserving Record Linkage, a problem in which the same real-world entity should be identified among databases from different dataholders, but without disclosing any additional information. In this paper, we investigate the potentials of Split Learning for Privacy-Preserving Record Matching, by introducing a novel training method through the utilization of Reference Sets, which are publicly available data corpora, showcasing minimal matching impact against a traditional centralized SVM-based technique.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。