arXiv:2509.02048cs.LGcs.AI2025-09

用几何曲率指导扰动,平衡数据隐私与可用性

Privacy-Utility Trade-off in Data Publication: A Bilevel Optimization Framework with Curvature-Guided Perturbation

  • 上层优化保质量,下层用曲率度量隐私风险
  • 在低曲率区扰动样本,显著提升抗成员推断攻击能力
  • 适合需要高隐私保障的数据发布场景

机器学习需数据训练,但直接共享原始数据易遭成员推断攻击(MIA)等隐私威胁。现有隐私保护方法如数据扰动、泛化和生成合成数据常降低数据准确性和多样性,影响下游任务性能。为此,我们提出一种双层优化框架:上层通过判别器引导生成高质量样本以维持数据效用;下层利用数据流形上的局部外在曲率作为个体对MIA脆弱性的量化指标,将样本扰动至低曲率区域,抑制易被攻击的特征组合。通过交替优化双目标,实现隐私与效用的协同平衡。大量实验表明,该方法不仅显著增强对MIA的抵抗能力,且在样本质量与多样性方面优于现有方法。

原文摘要 · Abstract (English)

Machine learning models require datasets for effective training, but directly sharing raw data poses significant privacy risk such as membership inference attacks (MIA). To mitigate the risk, privacy-preserving techniques such as data perturbation, generalization, and synthetic data generation are commonly utilized. However, these methods often degrade data accuracy, specificity, and diversity, limiting the performance of downstream tasks and thus reducing data utility. Therefore, striking an optimal balance between privacy preservation and data utility remains a critical challenge. To address this issue, we introduce a novel bilevel optimization framework for the publication of private datasets, where the upper-level task focuses on data utility and the lower-level task focuses on data privacy. In the upper-level task, a discriminator guides the generation process to ensure that perturbed latent variables are mapped to high-quality samples, maintaining fidelity for downstream tasks. In the lower-level task, our framework employs local extrinsic curvature on the data manifold as a quantitative measure of individual vulnerability to MIA, providing a geometric foundation for targeted privacy protection. By perturbing samples toward low-curvature regions, our method effectively suppresses distinctive feature combinations that are vulnerable to MIA. Through alternating optimization of both objectives, we achieve a synergistic balance between privacy and utility. Extensive experimental evaluations demonstrate that our method not only enhances resistance to MIA in downstream tasks but also surpasses existing methods in terms of sample quality and diversity.

隐私保护数据发布曲率分析双层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。