arXiv:2607.23987cs.LGcs.AI2026-07

在有限内存下优化流式联邦学习的数据准入与保留策略。

Adaptive Data Admission and Retention for Streaming Federated Learning

论文配图:Adaptive Data Admission and Retention for Streaming Federated Learning
图 1 · 摘自论文原文
  • 结合客户端留存规则与服务器端动态准入机制,自适应管理数据。
  • 实验显示该方法逼近理想基准,同时满足采样成本与存储约束。
  • 适合资源受限场景下的持续学习系统设计者参考。

我们研究了客户端内存受限的流式联邦学习问题,其中新生成的训练数据具有随时间变化的采样成本,需在时间上选择性地接纳与保留。考虑一个联合的服务器端准入与客户端内存管理框架,目标是在采样成本预算和缓冲区约束下最小化累积超额群体风险。首先,我们推导出一个学习误差界,明确捕捉了即时训练样本量、新样本增长以及重用不平衡对有效样本量的影响。基于此界获得的代理惩罚项,我们提出一种主动约束漂移加惩罚(ACDPP)策略,结合结构化的客户端K步保留规则、服务器端在线准入规则及随时间变化的矩形准入区域。通过一系列比较论证,借助辅助的恒定准入策略,将ACDPP的学习界与无成本的理想基准相连接。这给出了次线性遗憾和采样成本违规的显式保证,而缓冲区占用违规则通过离线选择保留时长加以控制。多个数据集上的实验表明,所提策略在满足采样成本与缓冲区约束的同时,仍能接近理想基准。

原文摘要 · Abstract (English)

We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and must be selectively admitted and retained over time. We consider a joint server-side admission and client-side memory-management framework with the objective of minimizing the cumulative excess population risk under a sampling-cost budget and buffer constraints. We first derive a learning-error bound that explicitly captures the effects of instantaneous training sample size, distinct-sample growth, and reuse imbalance through a characterization of the effective sample size. Through a surrogate penalty obtained from this bound, we develop an Active-Constraint Drift-Plus-Penalty (ACDPP) policy that combines a structured client-side $K$-step retention rule with a server-side online admission rule and a time-varying rectangular admission region. We further present a sequence of comparison arguments, via an auxiliary constant-admission policy, that connects the ACDPP learning bound to a costless oracle benchmark. This yields explicit guarantees in terms of sublinear regret and sampling-cost violation, while the buffer-occupancy violation is controlled through offline selection of the retention horizon. Experiments on multiple datasets demonstrate that the proposed policy remains close to the oracle benchmark while satisfying the sampling-cost and buffer constraints.

联邦学习流式学习内存管理优化策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。