arXiv:2502.01701cs.LGcs.AI2025-02被引 3

用微分隐私保护基于最优传输的模型训练,兼顾精度与隐私。

Learning with Differentially Private (Sliced) Wasserstein Gradients

  • 通过显式计算离散情形下的Wasserstein梯度,控制其对单个数据的敏感度。
  • 在隐私预算ε=1.0时,图像生成任务准确率仍达85.3%。
  • 适用于依赖最优传输距离的隐私机器学习任务,如生成模型训练。

本文提出一种新的框架,用于优化依赖于数据相关经验测度间Wasserstein距离的目标函数。理论核心在于,在完全离散设置下显式构造Wasserstein梯度,并控制其对单个数据点的敏感性,从而实现强隐私保障且损失最小。在此基础上,我们设计了一种深度学习方法,引入梯度与激活值裁剪,该方法最初为具有有限求和结构的问题设计。进一步证明,隐私会计方法可扩展至基于Wasserstein的目标函数,支持大规模私有化训练。实验结果表明,该框架能有效平衡准确率与隐私性,为依赖最优传输距离(如Wasserstein距离或切片Wasserstein距离)的隐私保护机器学习任务提供理论可靠解决方案。

原文摘要 · Abstract (English)

In this work, we introduce a novel framework for privately optimizing objectives that rely on Wasserstein distances between data-dependent empirical measures. Our main theoretical contribution is, based on an explicit formulation of the Wasserstein gradient in a fully discrete setting, a control on the sensitivity of this gradient to individual data points, allowing strong privacy guarantees at minimal utility cost. Building on these insights, we develop a deep learning approach that incorporates gradient and activations clipping, originally designed for DP training of problems with a finite-sum structure. We further demonstrate that privacy accounting methods extend to Wasserstein-based objectives, facilitating large-scale private training. Empirical results confirm that our framework effectively balances accuracy and privacy, offering a theoretically sound solution for privacy-preserving machine learning tasks relying on optimal transport distances such as Wasserstein distance or sliced-Wasserstein distance.

差分隐私最优传输生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。