分层披露信息,用人工审核提升隐私保护下的数据匹配精度。
Multi-Layer Privacy-Preserving Record Linkage with Clerical Review based on gradual information disclosure
- 多层主动学习机制,逐步披露信息并整合人工审核
- 仅需少量标注即可显著提升匹配准确率,隐私泄露风险低
- 适合对数据主权和隐私要求高的医疗、金融数据融合场景
隐私保护记录链接(PPRL)是敏感数据集成中的关键技术,其链接质量直接影响合并数据集及基于其上的机器学习应用的可用性。本文提出一种新型隐私保护协议,通过多层主动学习过程将人工审核融入PPRL中。不确定的匹配候选在多个层级上由人工与非人工判断源进行审查,以减少每条记录及总体的信息披露量。预测结果回传更新先前层级,从而提升未审核候选的链接性能。数据所有者可自主控制每条记录的披露程度,遵循最小必要原则与数据主权理念。在真实数据集上的实验表明,该方法在标注成本有限的情况下显著提升了链接质量,同时保持较低的隐私风险。
原文摘要 · Abstract (English)
Privacy-Preserving Record linkage (PPRL) is an essential component in data integration tasks of sensitive information. The linkage quality determines the usability of combined datasets and (machine learning) applications based on them. We present a novel privacy-preserving protocol that integrates clerical review in PPRL using a multi-layer active learning process. Uncertain match candidates are reviewed on several layers by human and non-human oracles to reduce the amount of disclosed information per record and in total. Predictions are propagated back to update previous layers, resulting in an improved linkage performance for non-reviewed candidates as well. The data owners remain in control of the amount of information they share for each record. Therefore, our approach follows need-to-know and data sovereignty principles. The experimental evaluation on real-world datasets shows considerable linkage quality improvements with limited labeling effort and privacy risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。