研究高维稀疏回归中隐私与准确性的权衡,揭示正则化如何提升隐私性。
Privacy-Accuracy Trade-offs in High-Dimensional LASSO under Perturbation Mechanisms
- 用近似消息传递法分析输出扰动和目标扰动两种隐私机制
- 强正则化可稳定估计器,改善隐私性;噪声过大反而降低稳定性
- 适合关注高维统计隐私的学者及数据安全应用开发者
我们研究高维情形下基于LASSO的隐私保护稀疏线性回归,分析两种常用差分隐私机制:输出扰动(向估计量添加噪声)和目标扰动(在损失函数中加入随机线性项)。利用近似消息传递(AMP)方法,在随机设计和隐私噪声条件下刻画这些估计量的典型行为。为量化隐私,采用平均情况下的KL散度,其具有邻近数据集可区分性的假设检验解释。分析表明,稀疏性在隐私-准确性权衡中起核心作用:更强的正则化能通过稳定估计量对抗单点数据变化来提升隐私性。进一步发现两种机制行为迥异:目标扰动中,增加噪声水平可能产生非单调效应,过量噪声会引发估计器失稳,导致对数据扰动更敏感。结果表明,AMP为高维稀疏模型中的隐私-准确性权衡分析提供了强大框架。
原文摘要 · Abstract (English)
We study privacy-preserving sparse linear regression in the high-dimensional regime, focusing on the LASSO estimator. We analyze two widely used mechanisms for differential privacy: output perturbation, which injects noise into the estimator, and objective perturbation, which adds a random linear term to the loss function. Using approximate message passing (AMP), we characterize the typical behavior of these estimators under random design and privacy noise. To quantify privacy, we adopt typical-case measures, including the on-average KL divergence, which admits a hypothesis-testing interpretation in terms of distinguishability between neighboring datasets. Our analysis reveals that sparsity plays a central role in shaping the privacy-accuracy trade-off: stronger regularization can improve privacy by stabilizing the estimator against single-point data changes. We further show that the two mechanisms exhibit qualitatively different behaviors. In particular, for objective perturbation, increasing the noise level can have non-monotonic effects, and excessive noise may destabilize the estimator, leading to increased sensitivity to data perturbations. Our results demonstrate that AMP provides a powerful framework for analyzing privacy-accuracy trade-offs in high-dimensional sparse models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。