arXiv:2409.07215stat.MLcs.CR2024-09被引 3

提出安全评估合并数据集价值的方法,保护隐私同时量化收益。

Is merging worth it? Securely evaluating the information gain for causal dataset acquisition

  • 用多方计算计算期望信息增益,不暴露原始数据
  • 合并可提升处理效应估计的重叠度与不确定性降低
  • 兼容差分隐私,适合医疗等敏感领域合作

跨机构合并数据集耗时耗力,尤其涉及私密信息时。数据持有方希望在不泄露敏感信息的前提下,预判哪些数据集合并最具价值。因果估计中,合并价值不仅取决于认知不确定性下降,还依赖于处理组与对照组重叠性的改善。为此,本文首次提出基于密码学安全的信息论方法,用于评估异质处理效应估计中合并数据集的预期信息增益(EIG)。通过多方计算实现对EIG的计算,确保原始数据不被披露。进一步证明该方法可与差分隐私(DP)结合,在满足任意隐私要求的同时,相较单独使用DP保持更高精度。据我们所知,这是首个专为因果估计设计的隐私保护数据集获取方法。在多种模拟与真实基准上验证了其有效性与可靠性。代码已公开:https://github.com/LucileTerminassian/causal_prospective_merge。

原文摘要 · Abstract (English)

Merging datasets across institutions is a lengthy and costly procedure, especially when it involves private information. Data hosts may therefore want to prospectively gauge which datasets are most beneficial to merge with, without revealing sensitive information. For causal estimation this is particularly challenging as the value of a merge depends not only on reduction in epistemic uncertainty but also on improvement in overlap. To address this challenge, we introduce the first cryptographically secure information-theoretic approach for quantifying the value of a merge in the context of heterogeneous treatment effect estimation. We do this by evaluating the Expected Information Gain (EIG) using multi-party computation to ensure that no raw data is revealed. We further demonstrate that our approach can be combined with differential privacy (DP) to meet arbitrary privacy requirements whilst preserving more accurate computation compared to DP alone. To the best of our knowledge, this work presents the first privacy-preserving method for dataset acquisition tailored to causal estimation. We demonstrate the effectiveness and reliability of our method on a range of simulated and realistic benchmarks. Code is publicly available: https://github.com/LucileTerminassian/causal_prospective_merge.

因果推断隐私计算数据融合差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。