解决高维在线学习中误差随批次增加而发散的问题
High-dimensional online learning via asynchronous decomposition: Non-divergent results, dynamic regularization, and beyond
- 用汇总统计构建代理得分函数,实现异步分解
- 误差边界不发散,保持稀疏性且渐进达到最优精度
- 适合需要长期稳定更新的实时高维数据场景
现有高维在线学习方法常面临误差界或每批样本量随数据批次增加而发散的问题。为此,我们提出一种异步分解框架,利用汇总统计构建当前批次学习的代理得分函数。该框架通过动态正则化迭代硬阈值算法实现,为稀疏在线优化提供高效且低内存的解决方案。我们提供了统一的理论分析,同时考虑流式计算误差与统计精度,证明所提估计器在所有批次中均保持非发散误差界和ℓ₀稀疏性。此外,随着批次累积,估计器自适应获得额外收益,渐近达到如同已知全历史数据及真实支持集时的最优精度。这一理论性质在广义线性模型实例中得到验证。
原文摘要 · Abstract (English)
Existing high-dimensional online learning methods often face the challenge that their error bounds, or per-batch sample sizes, diverge as the number of data batches increases. To address this issue, we propose an asynchronous decomposition framework that leverages summary statistics to construct a surrogate score function for current-batch learning. This framework is implemented via a dynamic-regularized iterative hard thresholding algorithm, providing a computationally and memory-efficient solution for sparse online optimization. We provide a unified theoretical analysis that accounts for both the streaming computational error and statistical accuracy, establishing that our estimator maintains non-divergent error bounds and $\ell_0$ sparsity across all batches. Furthermore, the proposed estimator adaptively achieves additional gains as batches accumulate, attaining the oracle accuracy as if the entire historical dataset were accessible and the true support were known. These theoretical properties are further illustrated through an example of the generalized linear model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。