arXiv:2601.12124cs.LGcs.AI2026-01被引 1

提出SynQP框架,评估生成数据的隐私风险与质量。

SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data

  • 用模拟敏感数据构建开放评测框架,保护原始数据隐私。
  • 差分隐私使身份泄露风险低于0.09监管阈值,优于非私密模型。
  • 新提出身份泄露风险指标,更准确衡量生成数据隐私性。

合成数据在医疗应用中引发隐私担忧,但缺乏公开的隐私评估框架阻碍了其推广。主要挑战在于难以获取敏感数据以构建基准数据集。为此,我们提出SynQP——一个基于模拟敏感数据的合成数据生成(SDG)隐私评估开源框架,确保原始数据保密。我们强调需采用能公平反映机器学习模型概率特性的隐私度量。作为示范,我们使用SynQP对CTGAN进行评估,并提出一种新的身份泄露风险指标,相比现有方法能更准确估计隐私风险。质量评估显示,非私密模型的机器学习效能达≥0.97;隐私评估(表II)表明,差分隐私持续降低身份泄露风险(SD-IDR)和成员推断攻击风险(SD-MIA),所有差分隐私增强模型均低于0.09的监管阈值。代码已开源。

原文摘要 · Abstract (English)

The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks, due to difficulties in acquiring sensitive data. To address this, we introduce SynQP, an open framework for benchmarking privacy in synthetic data generation (SDG) using simulated sensitive data, ensuring that original data remains confidential. We also highlight the need for privacy metrics that fairly account for the probabilistic nature of machine learning models. As a demonstration, we use SynQP to benchmark CTGAN and propose a new identity disclosure risk metric that offers a more accurate estimation of privacy risks compared to existing approaches. Our work provides a critical tool for improving the transparency and reliability of privacy evaluations, enabling safer use of synthetic data in health-related applications. % In our quality evaluations, non-private models achieved near-perfect machine-learning efficacy \(\ge0.97\). Our privacy assessments (Table II) reveal that DP consistently lowers both identity disclosure risk (SD-IDR) and membership-inference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold. Code available at https://github.com/CAN-SYNH/SynQP

合成数据隐私评估差分隐私医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。