arXiv:2602.17284cs.LG2026-02被引 6

提出高效计算随机采样隐私损耗的方法,提升差分隐私训练的效率与精度。

Efficient privacy loss accounting for subsampling and random allocation

  • 基于隐私损失分布(PLD)实现随机分配的快速隐私分析
  • 证明随机采样在差分隐私训练中隐私-效用平衡优于泊松采样
  • 新工具支持通用隐私损耗计算,无需针对噪声机制单独设计

我们研究了一种用户数据在 t 步中随机均匀选取 k 步进行使用的采样方案在差分隐私中的隐私放大特性。该方案已应用于差分隐私优化(Chua et al., 2024a;Choquette-Choo et al., 2025)和通信高效的高维私有聚合(Asi et al., 2026),相比标准泊松采样展现出更好的实用性。现有理论分析(Feldman & Shenfeld, 2025;Dong et al., 2025)虽接近泊松采样的界限,但仍存在两个主要缺陷:一是实际应用中隐私参数不够紧致,因分析过程包含近似;二是所计算的参数为保龄球棒或 Renyi 散度,使用时引入额外开销。本文证明,对任意差分隐私算法,随机分配的隐私损失分布(PLD)可高效计算。应用于高斯机制时,结果表明随机分配的隐私-效用权衡至少不劣于泊松子采样,尤其适合用于 DP-SGD 训练。为此,我们开发了基于 PLD 实现的新通用隐私损耗会计工具,使精确隐私分析拓展至子采样场景,不再依赖特定噪声机制的手动分析。

原文摘要 · Abstract (English)

We consider the privacy amplification properties of a sampling scheme in which a user's data isused in $k$ steps chosen randomly and uniformly from a sequence (or set) of $t$ steps. This sampling scheme has been recently applied in the context of differentially private optimization (Chua et al., 2024a; Choquette-Choo et al., 2025) and communication-efficient high-dimensional private aggregation (Asi et al., 2026), where it was shown to have utility advantages over the standard Poisson sampling. Theoretical analyses of this sampling scheme (Feldman & Shenfeld, 2025; Dong et al., 2025) lead to bounds that are close to those of Poisson sampling, yet still have two significant shortcomings. First, in many practical settings, the resulting privacy parameters are not tight due to the approximation steps in the analysis. Second, the computed parameters are either the hockey stick or Renyi divergence, both of which introduce overheads when used in privacy loss accounting. In this work, we demonstrate that the privacy loss distribution (PLD) of random allocation applied to any differentially private algorithm can be computed efficiently. When applied to the Gaussian mechanism, our results demonstrate that the privacy-utility trade-off for random allocation is at least as good as that of Poisson subsampling. In particular, random allocation is better suited for training via DP-SGD. To support these computations, our work develops new tools for general privacy loss accounting based on a notion of PLD realization. This notion allows us to extend accurate privacy loss accounting to subsampling which previously required manual noise-mechanism-specific analysis.

差分隐私隐私计算机器学习安全高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。