提出无批次限制的在线高维模型估计方法,提升精度与实用性。
Renewable Lasso without Batch-Number Constraints: A Gradient-Enhanced Approach

- 用梯度增强代理损失,仅依赖历史摘要实现在线更新。
- 在高维下保持误差可控,突破以往对批次数量的严格限制。
- 适用于分布式场景,客户端只需传梯度,适合隐私敏感数据。
我们研究高维广义线性模型在流式数据下的在线估计问题。针对非分布式场景,提出一种梯度增强的代理损失函数,仅利用历史摘要近似累积损失,改进了现有高维可再生估计方法,并消除了先前研究中的批次数量约束。进一步将该方法扩展至主-客户端架构下的分布式流数据,各站点仅需交换梯度向量而非完整数据。与直接套用Jordan等(2019)方法于二次代理损失不同,我们的调整方法无需客户端计算完整代理损失。在高维尺度下推导出非渐近误差界,不再依赖严格的批次数限制。线性和逻辑回归模型的模拟实验及真实数据应用均表明,该方法在准确性上优于现有可再生估计器。
原文摘要 · Abstract (English)
We study online estimation for high-dimensional generalized linear models with streaming data. First, for the non-distributed setting, we propose a gradient-enhanced surrogate loss that approximates the cumulative loss using only historical summaries, which modifies and improves upon the existing renewable estimation approach for the same model in the high-dimensional setting, and removes the batch-number constraint in previous studies. We then extend the method to distributed streaming data under the master-client architecture, where batches are partitioned across sites and only summaries (gradient vectors) are exchanged. Instead of directing applying the popular method of Jordan et al. (2019) to the surrogate quadratic loss, our adjusted approach does not require the clients to compute the full surrogate loss. We derive non-asymptotic error bounds under the high-dimensional scaling, without the stringent constraint on the number of batches in the previous studies. Simulation results under linear and logistic models, together with a real-data application, show improved accuracy over existing renewable estimators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。