arXiv:2507.19116cs.LGcs.AI2025-07

为开放图数据发布设计了带隐私保护的结构学习方法。

Graph Structure Learning with Privacy Guarantees for Open Graph Data

  • 在数据发布时注入结构化高斯噪声,直接保障隐私
  • 在严格隐私预算下仍能高精度恢复原始图结构
  • 适用于真实和合成数据,兼顾隐私与数据可用性

在数据发布方与使用方分离的场景中,公开图数据并保护个体隐私仍具挑战。尽管差分隐私(DP)提供严格保障,但多数方法仅在模型训练阶段应用,难以适用于开放数据场景。本文提出一种将高斯差分隐私(GDP)直接融入数据发布过程的隐私保护图结构学习框架。机制在发布前向原始数据注入结构化高斯噪声,提供正式的μ-GDP保证,并由此导出紧致的(ε, δ)-差分隐私边界。尽管存在隐私化带来的失真,我们证明可通过无偏惩罚似然法恢复原始稀疏逆协方差结构。进一步将框架扩展至离散数据,采用离散高斯噪声,同时保持隐私保证。在合成与真实世界数据集上的大量实验表明,该方法在严格隐私预算下仍实现优异的隐私-效用权衡,维持高图结构恢复准确率。结果建立了差分隐私理论与图模型隐私发布之间的正式联系。

原文摘要 · Abstract (English)

Publishing open graph data while preserving individual privacy remains challenging when data publishers and data users are distinct entities. Although differential privacy (DP) provides rigorous guarantees, most existing approaches enforce privacy during model training rather than at the data publishing stage. This limits the applicability to open-data scenarios. We propose a privacy-preserving graph structure learning framework that integrates Gaussian Differential Privacy (GDP) directly into the data release process. Our mechanism injects structured Gaussian noise into raw data prior to publication and provides formal $μ$-GDP guarantees, leading to tight $(\varepsilon, δ)$-differential privacy bounds. Despite the distortion introduced by privatization, we prove that the original sparse inverse covariance structure can be recovered through an unbiased penalized likelihood formulation. We further extend the framework to discrete data using discrete Gaussian noise while preserving privacy guarantees. Extensive experiments on synthetic and real-world datasets demonstrate strong privacy-utility trade-offs, maintaining high graph recovery accuracy under rigorous privacy budgets. Our results establish a formal connection between differential privacy theory and privacy-preserving data publishing for graphical models.

图学习差分隐私数据发布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。