arXiv:2508.16931cs.LGcs.AI2025-08中稿 · European Conferenc…

解决联邦学习中数据陈旧问题,平衡数据新鲜度与数量。

Degree of Staleness-Aware Data Updating in Federated Learning

  • 设计三参数调控的本地数据更新机制,协调数据新鲜度与数量。
  • 提出新指标DoS量化数据陈旧度,证明其与模型性能的定量关系。
  • 适用于对实时性要求高的场景,如工业监测、金融风控。

在高度时敏的联邦学习任务中,数据持续生成,数据陈旧严重影响模型性能。尽管已有工作通过调整本地数据更新频率或客户端选择策略来优化陈旧问题,但均未同时考虑数据陈旧度与数据量。本文提出DUFL(Data Updating in Federated Learning),一种激励机制,包含三个可调参数:服务器支付、过时数据保留率、客户端新鲜数据收集量,以协同优化本地数据的新鲜度与数量。为此,我们引入新的度量指标DoS(Degree of Staleness)来量化数据陈旧程度,并进行理论分析,揭示了DoS与模型性能之间的定量关系。我们将DUFL建模为具有动态约束的双阶段斯塔克尔伯格博弈,推导出客户端的闭式最优本地更新策略及服务器的近似最优策略。在真实数据集上的实验结果表明,该方法显著提升模型性能。

原文摘要 · Abstract (English)

Handling data staleness remains a significant challenge in federated learning with highly time-sensitive tasks, where data is generated continuously and data staleness largely affects model performance. Although recent works attempt to optimize data staleness by determining local data update frequency or client selection strategy, none of them explore taking both data staleness and data volume into consideration. In this paper, we propose DUFL(Data Updating in Federated Learning), an incentive mechanism featuring an innovative local data update scheme manipulated by three knobs: the server's payment, outdated data conservation rate, and clients' fresh data collection volume, to coordinate staleness and volume of local data for best utilities. To this end, we introduce a novel metric called DoS(the Degree of Staleness) to quantify data staleness and conduct a theoretic analysis illustrating the quantitative relationship between DoS and model performance. We model DUFL as a two-stage Stackelberg game with dynamic constraint, deriving the optimal local data update strategy for each client in closed-form and the approximately optimal strategy for the server. Experimental results on real-world datasets demonstrate the significant performance of our approach.

联邦学习数据陈旧激励机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。