arXiv:2605.16219cs.LGstat.ML2026-05

差分隐私让尾部风险学习的有效样本量从n降至nτ,带来额外隐私代价。

The Privacy Price of Tail-Risk Learning: Effective Tail Sample Size in Differentially Private CVaR Optimization

  • 用有效尾部样本量nτ替代n,揭示隐私对尾部风险学习的影响
  • 在纯差分隐私下,尾部风险误差率达到Θ(1/(εnτ))
  • 适用于关注隐私保护的金融风控、极端事件建模的研究者

差分隐私会改变CVaR学习的有效样本量。对于尾部质量τ,隐私相关的样本量不是n,而是nτ;等价地,有效私有尾部样本量为εnτ。私有CVaR超额风险可分解为普通尾部风险统计误差与隐私代价。该分解在标量估计和有限类情况下是完整的:标量估计的误差率为Θ(B min{1, (nτ)^{-1/2} + (εnτ)^{-1}}),有限类规模为M时的误差率为Θ(B min{1, √(log(2M)/(nτ)) + log(2M)/(εnτ)})。这些完整速率在纯差分隐私下成立,其下界在近似差分隐私的小δ范围内也成立。对于凸Lipschitz学习,模块化上下界归约表明,专用于CVaR的隐私项必然以1/(εnτ)缩放,维度依赖继承自私有随机凸优化。这些结果将私有CVaR学习中的核心难题识别为在Θ(nτ)个有信息的尾部记录上进行私有学习。

原文摘要 · Abstract (English)

Differential privacy changes the effective sample size governing CVaR learning. For tail mass $τ$, the privacy-relevant sample size is not $n$, but $nτ$; equivalently, the effective private tail sample size is $εnτ$. Private CVaR excess risk decomposes into ordinary tail-risk statistical error and a privacy price. This decomposition is complete for scalar estimation and finite classes: scalar estimation has rate $Θ(B \min\{1,(nτ)^{-1/2}+(εnτ)^{-1}\})$, and finite classes of size $M$ have rate $Θ(B \min\{1,\sqrt{\log(2M)/(nτ)}+\log(2M)/(εnτ)\})$. These complete rates hold under pure DP, and their lower bounds extend to approximate DP in the stated small-$δ$ regimes. For convex Lipschitz learning, modular upper and lower reductions show that the CVaR-specific privacy term necessarily scales as $1/(εnτ)$, with dimension dependence inherited from private stochastic convex optimization. Together, these results identify ordinary private learning on $Θ(nτ)$ informative tail records as the canonical hard subproblem inside private CVaR learning.

差分隐私风险建模尾部分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。