提出端到端协作生成合成数据框架,支持隐私保护下的多机构数据共享。
End to End Collaborative Synthetic Data Generation
- 通过安全多方计算实现隐私保护的数据预处理与合成
- 在白血病基因组数据上验证可生成高质量合成数据
- 适合医疗等敏感领域跨机构协作研究
AI的成功依赖于训练数据的可用性。虽然某些情况下单一数据持有方已具备足够数据,但更多情形下需多个机构协作以达成有意义的AI研究所需的数据规模,例如罕见疾病研究中各临床中心仅掌握少量患者数据。近年来联邦合成数据生成算法为隐私保护下的协作数据共享迈出关键一步。然而现有方法仅关注合成器训练,假设数据已预处理完毕且合成结果可一次性交付,无需超参数调优。本文提出一种端到端协作框架,涵盖隐私保护的数据预处理与评估环节。我们基于安全多方计算(MPC)协议实现该框架,并在白血病基因组数据的隐私保护合成发布场景中进行评估。
原文摘要 · Abstract (English)
The success of AI is based on the availability of data to train models. While in some cases a single data custodian may have sufficient data to enable AI, often multiple custodians need to collaborate to reach a cumulative size required for meaningful AI research. The latter is, for example, often the case for rare diseases, with each clinical site having data for only a small number of patients. Recent algorithms for federated synthetic data generation are an important step towards collaborative, privacy-preserving data sharing. Existing techniques, however, focus exclusively on synthesizer training, assuming that the training data is already preprocessed and that the desired synthetic data can be delivered in one shot, without any hyperparameter tuning. In this paper, we propose an end-to-end collaborative framework for publishing of synthetic data that accounts for privacy-preserving preprocessing as well as evaluation. We instantiate this framework with Secure Multiparty Computation (MPC) protocols and evaluate it in a use case for privacy-preserving publishing of synthetic genomic data for leukemia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。