arXiv:2502.18578cs.LGcs.AI2025-02被引 2

首个面向线性回归的差分隐私筛选规则,解决隐私保护下的特征筛选难题。

Differentially Private Iterative Screening Rules for Linear Regression

  • 基于差分隐私设计新型特征筛选规则,实现隐私保护下的稀疏回归。
  • 原版规则过度筛选特征,削弱后显著减少误删,提升模型性能。
  • 适合关注隐私计算与高维数据建模的研究者使用。

L₁正则化线性模型是数据科学中简单而有效的工具。过去十年中,特征筛选规则因能有效生成稀疏回归系数而日益流行。然而,随着对数据隐私保护的需求增加,目前尚无可用于此类模型的差分隐私筛选规则。本文首次提出针对线性回归的差分隐私筛选规则。研究发现,原始版本的规则过于严格,导致过多系数被错误剔除。通过弱化实现方式,可有效缓解过度筛选问题,显著提升模型性能。

原文摘要 · Abstract (English)

Linear $L_1$-regularized models have remained one of the simplest and most effective tools in data science. Over the past decade, screening rules have risen in popularity as a way to eliminate features when producing the sparse regression weights of $L_1$ models. However, despite the increasing need of privacy-preserving models for data analysis, to the best of our knowledge, no differentially private screening rule exists. In this paper, we develop the first private screening rule for linear regression. We initially find that this screening rule is too strong: it screens too many coefficients as a result of the private screening step. However, a weakened implementation of private screening reduces overscreening and improves performance.

差分隐私特征筛选线性回归稀疏学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。