针对流式异构数据,提出联邦在线学习方法,兼顾隐私与高效更新。
Federated Online Learning for Heterogeneous Multisource Streaming Data
- 为每个数据源构建个性化模型,通过分组假设捕捉相似性提升性能。
- 仅需传递前批次统计量,存储开销低,且不共享原始数据保障隐私。
- 理论证明模型一致性和变量选择最优性,适用于金融与网络日志等场景。
联邦学习已成为隐私约束下分布式多源数据分析的重要范式。现有方法多聚焦于静态数据集,而现实应用中数据持续流入形成流式数据,带来存储与算法设计新挑战,尤其在高维场景下。本文提出一种用于分布式多源流式数据的联邦在线学习(FOL)方法。为应对异质性,为每个数据源构建个性化模型,并引入新颖的“子群”假设以捕捉潜在相似性,从而提升模型性能。采用惩罚可更新估计方法与高效的近端梯度下降进行训练。所提方法同时符合联邦学习与在线学习框架:原始数据不跨源交换,确保数据隐私;模型更新仅需前批次数据的汇总统计量,显著降低存储需求。理论上,建立了模型估计、变量选择及子群结构恢复的一致性性质,证明了最优统计效率。模拟实验验证了方法有效性。应用于金融信贷数据与网页日志数据时,亦展现出优越的预测性能,分析结果提供实用洞见。
原文摘要 · Abstract (English)
Federated learning has emerged as an essential paradigm for distributed multi-source data analysis under privacy concerns. Most existing federated learning methods focus on the ``static" datasets. However, in many real-world applications, data arrive continuously over time, forming streaming datasets. This introduces additional challenges for data storage and algorithm design, particularly under high-dimensional settings. In this paper, we propose a federated online learning (FOL) method for distributed multi-source streaming data analysis. To account for heterogeneity, a personalized model is constructed for each data source, and a novel ``subgroup" assumption is employed to capture potential similarities, thereby enhancing model performance. We adopt the penalized renewable estimation method and the efficient proximal gradient descent for model training. The proposed method aligns with both federated and online learning frameworks: raw data are not exchanged among sources, ensuring data privacy, and only summary statistics of previous data batches are required for model updates, significantly reducing storage demands. Theoretically, we establish the consistency properties for model estimation, variable selection, and subgroup structure recovery, demonstrating optimal statistical efficiency. Simulations illustrate the effectiveness of the proposed method. Furthermore, when applied to the financial lending data and the web log data, the proposed method also exhibits advantageous prediction performance. Results of the analysis also provide some practical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。