arXiv:2605.25916cs.LGcs.DC2026-05

用强化学习统一优化边缘设备的训练与推理,提升效率与精度。

Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning

论文配图:Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning
图 1 · 摘自论文原文
  • 设计双队列机制连接推理请求与训练数据,融合数据与模型新鲜度。
  • 在多种配置下显著降低延迟和能耗,同时保持高推理准确率。
  • 适合资源受限的边缘智能场景,如物联网与实时计算系统。

联邦边缘学习(FEEL)作为一种新兴范式,通过在边缘设备间协同训练模型,实现边缘智能(EI),同时保护数据隐私。本文提出一种在线优化框架,联合管理资源受限边缘设备上的联邦训练与推理。引入受双队列启发的转换机制,连接推理请求与训练数据,并将数据与模型的新鲜度纳入准确率建模,以捕捉真实环境中的时序动态。为在最大化推理准确率的同时最小化延迟与能耗,对边缘设备的模式选择、通信与计算资源分配进行联合优化。将该问题建模为多目标优化问题,其为NP难且受在线设置进一步复杂化。为此,将其转化为多目标马尔可夫决策过程(MOMDP),并提出约束多目标近端策略优化(C-MOPPO)算法。具体地,C-MOPPO首先学习一组具有不同目标偏好的策略,再通过约束策略优化扩展帕累托前沿,获得高质量、密集的解。大量实验表明,C-MOPPO在各类系统配置下均实现了目标间的良好权衡,显著优于基线方法。

原文摘要 · Abstract (English)

Federated edge learning (FEEL) has recently emerged as a promising paradigm for achieving edge intelligence (EI) via enabling collaborative model training across edge devices while protecting data privacy. In this paper, we put forth an online optimization framework that jointly manages federated training and inference on resource-constrained edge devices. We introduce a tandem-queue-inspired conversion mechanism that bridges inference requests and training data, and further incorporate both data and model freshness into the accuracy formulation to capture temporal dynamics in real-world environments. To maximize inference accuracy while minimizing latency and energy consumption, the mode selections, communication, and computation resource allocations of edge devices are jointly optimized. We formulate this optimization as a multi-objective optimization problem, which is NP-hard and further complicated by the online setting. To address these challenges, we transform the problem into a multi-objective Markov decision process (MOMDP) and develop a \underline{c}onstrained \underline{m}ulti-\underline{o}bjective \underline{p}roximal \underline{p}olicy \underline{o}ptimization (C-MOPPO) algorithm. Specifically, C-MOPPO first learns a set of policies with different preferences across three objectives, then leverages constrained policy optimization to enrich the Pareto front and obtain high-quality, dense solutions. Extensive experiments demonstrate that C-MOPPO achieves well-balanced trade-offs among objectives and significantly outperforms baselines under various system configurations.

边缘计算联邦学习强化学习多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。