arXiv:2607.29675math.STcs.LG2026-07

提出隐私保护的密度模态学习方法,实现高精度模式恢复与隐私保障。

Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering

论文配图:Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering
图 1 · 摘自论文原文
  • 基于均值迁移思想,结合高阶核函数与梯度裁剪加噪,实现私密梯度上升
  • 在多变量分布下证明模式可高概率恢复,误差率逼近最优水平
  • 适用于回归与聚类任务,隐私-效用权衡优于主流基线方法

密度模态为多模态分布提供了局部且可解释的总结,但在严格的差分隐私约束下其估计仍鲜有研究。本文在局部光滑性、曲率和分离条件下,研究多变量分布下密度模态的差分隐私恢复问题。提出DP-GRAMS方法,受均值迁移启发,在差分隐私评分估计基础上进行带噪声的梯度上升。假设密度局部属于Hölder类且光滑参数β>2,采用偏置减少的高阶核函数,并通过梯度裁剪和校准高斯噪声实现梯度上升步骤的隐私保护。私有初始化方案结合密度感知效用与抑制规则,利用k∼M log n次采样于公开h_DAP网格,抑制半径ρ_init∼(log n)^(-1/d),通过逐步抑制竞争区域的局部邻域,实现模态盆地的高概率覆盖;多次启动间的相关噪声使联合发布仅需单一(ε,δ)差分隐私保证。理论证明所有总体模态均可高概率恢复,并建立渐近误差率:O((log n/n)^{2(β−1)/(d+2β)}) + O((polylog(n,δ)/(n²ε²))^{(β−1)/(d+β)})。同时给出私密模态估计的极小极大下界,表明所提估计器几乎最优,仅差对数因子。进一步提出两种自然扩展:DP-PMS(私密模态回归)与DP-GRAMS-C(聚类流程)。合成与真实数据上的大量实验显示,相比常见基线,该方法在隐私-效用权衡上表现更优。

原文摘要 · Abstract (English)

Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored. We study differentially private recovery of density modes for multivariate distributions under local smoothness, curvature, and separation conditions. We propose DP-GRAMS, a mean-shift inspired method that performs noisy ascent on a differentially private score estimator. Assuming the density belongs locally to a Hölder class with smoothness parameter $β> 2$, our score estimator uses bias-reducing higher-order kernels, and then enforces privacy in the gradient ascent steps via gradient clipping and calibrated Gaussian noise. A private initialization scheme combines a density-aware utility with a suppression rule and, with $k\asymp M\log n$ draws over a public $h_{\mathrm{DAP}}$-grid and suppression radius $ρ_{\mathrm{init}}\asymp (\log n)^{-1/d}$, achieves high-probability coverage of the modal basins by successively suppressing selected local neighborhoods in competitive regions, while correlated noise across multiple starts enables joint release under a single $(\varepsilon,δ)$-differential privacy guarantee. We prove that all population modes are recovered with high probability and establish asymptotic error rates of the form $O\!\left((\tfrac{\log n}{n})^{\frac{2(β-1)}{d+2β}}\right) + O\!\left((\tfrac{\mathrm{polylog}(n,δ)}{n^2\varepsilon^2})^{\frac{β-1}{d+β}}\right)$. We also provide minimax lower bounds for private mode estimation, and show that our estimators are nearly optimal, up to a logarithmic factor in the MSE. We present two natural extensions: DP-PMS, a private modal-regression method, and DP-GRAMS-C, a clustering pipeline. Extensive experiments on synthetic and real data demonstrate favorable privacy-utility trade-offs relative to common baselines.

差分隐私密度估计聚类非参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。