arXiv:2409.17315cs.LGcs.AI2024-09被引 2

用知识图谱增强生成模型,让隐私保护数据更真实可信。

KIPPS: Knowledge infusion in Privacy Preserving Synthetic Data Generation

  • 将知识图谱中的领域规则注入生成模型,指导数据生成
  • 在医疗与网络安全数据上实现高隐私性与高准确率平衡
  • 适合需要合规生成数据的敏感领域研究者使用

将差分隐私等隐私保护技术融入合成数据生成,可提供可证明的隐私保障。然而,在网络安全、医疗等关键领域,生成式深度学习模型在处理离散、非高斯且受约束的特征时面临挑战,尤其当训练数据有限且多样性不足时,模型易重复敏感特征,造成隐私泄露风险。同时,模型难以理解特定领域的属性约束,导致生成不现实数据,影响下游任务准确性。为此,本文提出新模型KIPPS,通过知识图谱注入领域与监管知识,为生成模型提供额外上下文信息,并在训练中强制执行领域约束。该框架显著提升模型生成真实、合规合成数据的能力。在真实世界医疗与网络安全数据集上的实验表明,该方法在保持高隐私韧性的同时,优于基准方法,有效平衡了隐私保护与数据准确性。

原文摘要 · Abstract (English)

The integration of privacy measures, including differential privacy techniques, ensures a provable privacy guarantee for the synthetic data. However, challenges arise for Generative Deep Learning models when tasked with generating realistic data, especially in critical domains such as Cybersecurity and Healthcare. Generative Models optimized for continuous data struggle to model discrete and non-Gaussian features that have domain constraints. Challenges increase when the training datasets are limited and not diverse. In such cases, generative models create synthetic data that repeats sensitive features, which is a privacy risk. Moreover, generative models face difficulties comprehending attribute constraints in specialized domains. This leads to the generation of unrealistic data that impacts downstream accuracy. To address these issues, this paper proposes a novel model, KIPPS, that infuses Domain and Regulatory Knowledge from Knowledge Graphs into Generative Deep Learning models for enhanced Privacy Preserving Synthetic data generation. The novel framework augments the training of generative models with supplementary context about attribute values and enforces domain constraints during training. This added guidance enhances the model's capacity to generate realistic and domain-compliant synthetic data. The proposed model is evaluated on real-world datasets, specifically in the domains of Cybersecurity and Healthcare, where domain constraints and rules add to the complexity of the data. Our experiments evaluate the privacy resilience and downstream accuracy of the model against benchmark methods, demonstrating its effectiveness in addressing the balance between privacy preservation and data accuracy in complex domains.

隐私生成知识图谱合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。