arXiv:2512.08869cs.LGcs.AI2025-12被引 7

用规则感知的生成模型,让合成数据既真实又隐私安全。

Differentially Private Synthetic Data Generation Using Context-Aware GANs

  • 用约束矩阵融合显性和隐性规则,指导生成过程。
  • 在医疗、金融等领域生成的数据更符合实际且保护隐私。
  • 适合需合规与真实性的高敏感数据应用。

大数据广泛应用引发严重隐私担忧,尤其在医疗、金融等领域。如GDPR和HIPAA等法规对数据处理要求严格,难以兼顾数据分析需求与隐私保护。合成数据提供了一种解决方案:生成反映真实模式但不暴露敏感信息的人工数据集。然而,传统方法常无法捕捉数据中隐含的复杂规则(如特定疾病禁用某些药物、避免有害药物相互作用),导致生成数据虽有表面模式却缺乏现实合理性。为解决此问题,本文提出ContextGAN——一种结合领域规则的差分隐私生成对抗网络。通过约束矩阵编码显性与隐性知识,约束型判别器评估生成数据是否满足领域规则,同时差分隐私保障原始数据敏感信息不泄露。我们在医疗、安全、金融三个领域验证了ContextGAN,结果表明其生成的数据在保持隐私的同时显著提升真实性与实用性,特别适用于需遵守显性模式与隐性规则且具备强隐私保障的应用场景。

原文摘要 · Abstract (English)

The widespread use of big data across sectors has raised major privacy concerns, especially when sensitive information is shared or analyzed. Regulations such as GDPR and HIPAA impose strict controls on data handling, making it difficult to balance the need for insights with privacy requirements. Synthetic data offers a promising solution by creating artificial datasets that reflect real patterns without exposing sensitive information. However, traditional synthetic data methods often fail to capture complex, implicit rules that link different elements of the data and are essential in domains like healthcare. They may reproduce explicit patterns but overlook domain-specific constraints that are not directly stated yet crucial for realism and utility. For example, prescription guidelines that restrict certain medications for specific conditions or prevent harmful drug interactions may not appear explicitly in the original data. Synthetic data generated without these implicit rules can lead to medically inappropriate or unrealistic profiles. To address this gap, we propose ContextGAN, a Context-Aware Differentially Private Generative Adversarial Network that integrates domain-specific rules through a constraint matrix encoding both explicit and implicit knowledge. The constraint-aware discriminator evaluates synthetic data against these rules to ensure adherence to domain constraints, while differential privacy protects sensitive details from the original data. We validate ContextGAN across healthcare, security, and finance, showing that it produces high-quality synthetic data that respects domain rules and preserves privacy. Our results demonstrate that ContextGAN improves realism and utility by enforcing domain constraints, making it suitable for applications that require compliance with both explicit patterns and implicit rules under strict privacy guarantees.

合成数据差分隐私生成模型医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。