arXiv:2411.00879cs.DBcs.LG2024-11被引 2

解决数据协作中重复主体的合成难题,提升隐私保护下的数据共享效果。

DEREC-SIMPRO: unlock Language Model benefits to advance Synthesis in Data Clean Room

  • 提出DEREC预处理三步法,增强多表合成模型对重复主体的适应性。
  • 使用SIMPRO评估指标,发现多表合成效果优于单表模型。
  • 适合关注数据隐私与跨表协作的研究者和工业应用团队。

通过数据洁净室进行数据协作虽具价值,但引发隐私担忧,可通过合成数据和多表生成器缓解。然而现有多表生成器在两表中存在重复主体时性能显著下降,这一问题普遍存在且亟待解决。为此,本文提出DEREC三步预处理流程,提升多表生成器对重复主体场景的泛化能力;同时引入SIMPRO三维度评估指标,结合条件分布分析与大规模同步假设检验,从列级与表级全面评估合成数据的真实性。实验表明,采用DEREC可显著提升合成数据保真度,多表生成器在协作场景中优于单表模型。DEREC-SIMPRO整体方案为通用数据协作提供了稳健解决方案,推动更高效、数据驱动的社会发展。

原文摘要 · Abstract (English)

Data collaboration via Data Clean Room offers value but raises privacy concerns, which can be addressed through synthetic data and multi-table synthesizers. Common multi-table synthesizers fail to perform when subjects occur repeatedly in both tables. This is an urgent yet unresolved problem, since having both tables with repeating subjects is common. To improve performance in this scenario, we present the DEREC 3-step pre-processing pipeline to generalize adaptability of multi-table synthesizers. We also introduce the SIMPRO 3-aspect evaluation metrics, which leverage conditional distribution and large-scale simultaneous hypothesis testing to provide comprehensive feedback on synthetic data fidelity at both column and table levels. Results show that using DEREC improves fidelity, and multi-table synthesizers outperform single-table counterparts in collaboration settings. Together, the DEREC-SIMPRO pipeline offers a robust solution for generalizing data collaboration, promoting a more efficient, data-driven society.

数据合成隐私保护多表生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。