提出高效通信的稀疏回归算法,避免传输原始数据
Communication-Efficient l_0 Penalized Least Square
- 基于分布式系统优化通信,通过主动集更新减少数据传输
- 在不传原始数据情况下达到全局估计器的统计精度
- 适合大规模隐私敏感数据的分布式建模场景
本文提出一种针对高维稀疏线性回归模型的通信高效惩罚回归算法。该方法结合名为CESDAR的优化分布式通信算法,基于增强支持检测与根查找算法,在多机环境下计算并更新活跃变量集,并引入通信高效的代理似然框架,在活跃集上近似全样本最优解,从而避免原始数据传输,提升算法执行速度并大幅降低通信开销。该方法在保持与全局估计器相同统计精度的同时,显著增强了隐私保护与数据安全性。此外,本文还探讨了CESDAR的扩展版本和自适应版本,分别用于提升算法速度与优化参数选择。仿真与真实数据基准实验均验证了该算法的效率与准确性。
原文摘要 · Abstract (English)
In this paper, we propose a communication-efficient penalized regression algorithm for high-dimensional sparse linear regression models with massive data. This approach incorporates an optimized distributed system communication algorithm, named CESDAR algorithm, based on the Enhanced Support Detection and Root finding algorithm. The CESDAR algorithm leverages data distributed across multiple machines to compute and update the active set and introduces the communication-efficient surrogate likelihood framework to approximate the optimal solution for the full sample on the active set, resulting in the avoidance of raw data transmission, which enhances privacy and data security, while significantly improving algorithm execution speed and substantially reducing communication costs. Notably, this approach achieves the same statistical accuracy as the global estimator. Furthermore, this paper explores an extended version of CESDAR and an adaptive version of CESDAR to enhance algorithmic speed and optimize parameter selection, respectively. Simulations and real data benchmarks experiments demonstrate the efficiency and accuracy of the CESDAR algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。