arXiv:2409.16832cs.LGcs.NI2024-09被引 2

提出异步分式多智能体强化学习,降低移动边缘计算中的信息年龄。

Asynchronous Fractional Multi-Agent Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing

  • 基于分式强化学习框架,联合优化任务生成与卸载策略
  • 在真实场景下平均信息年龄降低50.6%(相比最优基线)
  • 无需实时系统状态即可异步决策,适合动态边缘计算场景

在如网络物理系统(CPS)等实时网络应用中,信息年龄(AoI)是衡量时效性的关键指标。为满足智能制造等高算力需求,移动边缘计算(MEC)可有效优化计算并降低AoI。本文研究计算密集型更新的时效性,联合优化任务生成时机与卸载位置以最小化AoI。考虑边缘负载动态,构建最小化期望时间平均AoI的任务调度问题。由于AoI带来的分式目标及半马尔可夫博弈(SMG)中的异步决策,求解极具挑战。为此,提出分式强化学习(RL)框架:先建立分式单智能体RL并证明其线性收敛;再扩展为分式多智能体框架,结合Dinkelbach方法,证明其等价于非精确牛顿法,并给出达到纳什均衡线性收敛的条件。针对异步决策难题,设计异步无模型分式多智能体算法,各移动设备无需了解其他设备状态或系统动态即可自主决策。实验表明,相比最优基线,平均AoI降低达50.6%。

原文摘要 · Abstract (English)

In the realm of emerging real-time networked applications such as cyber-physical systems (CPS), the Age of Information (AoI) has emerged as a pivotal metric for evaluating timeliness. To meet the high computational demands, such as those in smart manufacturing within CPS, mobile edge computing (MEC) presents a promising solution for optimizing computing and reducing AoI. In this work, we study the timeliness of compute-intensive updates and explore jointly optimizing the task updating (when to generate a task) and offloading (where to process a task) policies to minimize AoI. Specifically, we consider edge load dynamics and formulate a task scheduling problem to minimize the expected time-average AoI. Solving this problem is challenging due to the fractional objective introduced by AoI and the asynchronous decision-making of the semi-Markov game (SMG). To this end, we propose a fractional reinforcement learning (RL) framework. We begin by introducing a fractional single-agent RL framework and establish its linear convergence rate. Building on this, we develop a fractional multi-agent RL framework, extend Dinkelbach's method, and demonstrate its equivalence to the inexact Newton's method. Furthermore, we provide the conditions under which the framework achieves linear convergence to the Nash equilibrium (NE). To tackle the challenge of asynchronous decision-making in the SMG, we further design an asynchronous model-free fractional multi-agent RL algorithm, where each mobile device can determine the task updating and offloading decisions without knowing the real-time system dynamics and decisions of other devices. Experimental results show that when compared with the best existing baseline algorithm, our proposed algorithm reduces the average AoI by up to 50.6%.

边缘计算强化学习信息年龄多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。