让机构在不泄露隐私的前提下,安全评估外部数据价值,促进数据合作。
Privacy-Preserving Dataset Combination
- 采用私有计算技术,在不暴露数据的情况下评估源数据与候选数据的差异性。
- 真实数据测试显示,评估结果与非私有方法相关性超90%,准确识别有益合作。
- 适合医疗、金融等敏感领域中小机构,解决数据共享中的隐私与效用矛盾。
高质量、多样化的数据集对机器学习模型性能至关重要,但隐私顾虑和竞争利益限制了数据共享,尤其在医疗等受监管领域。小型机构因缺乏资源购买数据或达成有利共享协议而处于劣势,原因在于无法安全评估外部数据的实用性。为同时解决隐私与不确定性问题,本文提出{ SecureKL},首个无需泄露隐私即可进行数据集间评估的安全协议,适用于数据共享前的预评估阶段。该方法在内部通过私有计算执行数据集差异度量,无需依赖下游模型。在真实数据上,{ SecureKL}实现与非私有方法>90%的相关性,并成功识别出跨医院重症监护室死亡率预测及跨州收入预测等高度异构场景下的有效数据合作。结果表明,安全计算能最大化数据利用效率,优于泄露信息的非隐私感知评估方式。
原文摘要 · Abstract (English)
Access to diverse, high-quality datasets is crucial for machine learning model performance, yet data sharing remains limited by privacy concerns and competitive interests, particularly in regulated domains like healthcare. This dynamic especially disadvantages smaller organizations that lack resources to purchase data or negotiate favorable sharing agreements, due to the inability to \emph{privately} assess external data's utility. To resolve privacy and uncertainty tensions simultaneously, we introduce {\SecureKL}, the first secure protocol for dataset-to-dataset evaluations with zero privacy leakage, designed to be applied preceding data sharing. {\SecureKL} evaluates a source dataset against candidates, performing dataset divergence metrics internally with private computations, all without assuming downstream models. On real-world data, {\SecureKL} achieves high consistency ($>90\%$ correlation with non-private counterparts) and successfully identifies beneficial data collaborations in highly-heterogeneous domains (ICU mortality prediction across hospitals and income prediction across states). Our results highlight that secure computation maximizes data utilization, outperforming privacy-agnostic utility assessments that leak information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。