针对流式数据的分布式优化,提出时间加权新视角,提升动态环境下的学习效果。
Distributed Optimization with Streaming Data: A Temporal Weighting Perspective

- 将全局目标设为随时间加权的损失均值,建模动态学习过程
- 理论证明在常步长下存在非零偏差底限,受网络连通性影响
- 不同加权策略(如指数衰减、滑动窗口)影响跟踪误差的长期表现
优化理论是智能决策的常用工具。传统优化处理固定、时不变的目标函数,但许多现代应用在动态环境中运行,数据持续到达,学习目标随时间演变,且面临分布式数据与通信限制。为此,我们从流式数据出发,通过结构化的时变公式研究分布式优化,其中全局目标是网络中各处观测损失的时间加权平均。我们分析多轮次分布式一阶方法,包括分布式梯度下降。对于强凸且光滑的损失函数,通过压缩映射视角建立了欧氏范数的跟踪误差保证。结果表明,跟踪误差可分解为固定点跟踪分量和由去中心化与数据异质性引起的偏差项。我们将分析特化到均匀加权、指数折扣加权及其有限记忆的滑动窗口版本。边界明确刻画了时间加权规则、每步迭代预算、步长和网络连通性的作用。均匀加权下固定点跟踪贡献为 $\mathcal O(1/t)$,而折扣与滑动窗口策略通常导致由折扣因子和有效记忆决定的非零跟踪底限。在所有情况下,常步长下去中心化都会引入额外的非零偏差底限。数值实验验证了预测趋势。
原文摘要 · Abstract (English)
Optimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant objective functions, many modern applications operate in dynamic environments where data arrive sequentially, and the learning objective evolves over time, often under decentralized data and communication constraints. Motivated by these trends, we study decentralized optimization from streaming data through a structured time-varying formulation in which the global objective is a temporally weighted average of losses observed across the network. We analyze multi-iteration decentralized first-order methods, including decentralized gradient descent. For strongly convex and smooth losses, we develop guarantees for the Euclidean-norm \emph{tracking error} through a contraction-mapping viewpoint. The resulting bounds decompose the tracking error into a fixed-point tracking component and a bias term induced by decentralization and data heterogeneity. We specialize our analysis to uniform and exponentially discounted weights, as well as their finite-memory \emph{windowed} counterparts. The bounds explicitly characterize the roles of the temporal weighting rule, per-step iteration budget, step size, and network connectivity. Uniform weighting yields a vanishing fixed-point tracking contribution of order $\mathcal O(1/t)$, whereas discounted and windowed strategies generally induce non-vanishing tracking floors governed by the discount factor and effective memory, respectively. In all cases, decentralization induces an additional non-zero bias floor under a constant step size. Numerical experiments illustrate the predicted trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。