多智能体强化学习优化无线充电网络寿命,支持高效协同充电。
Multi-agent reinforcement learning strategy to maximize the lifetime of Wireless Rechargeable
- 设计去中心化部分可观测半马尔可夫决策模型,实现移动充电器协作。
- 单次充电可覆盖多个传感器,显著提升充电效率与网络寿命。
- 无需重新训练即可适配不同网络,适合大规模无线充电场景。
本文提出一种通用的多移动充电器充电框架,旨在最大化大规模无线可充电传感网络(WRSNs)的生命周期,同时保障目标覆盖与连通性。引入多点充电模型,使移动充电器(MC)可在每个充电位置同时为多个传感器供电,提升充电效率。提出一种去中心化的部分可观测半马尔可夫决策过程(Dec POSMDP)模型,促进移动充电器间协作,并基于实时网络信息识别最优充电位置。此外,该方法支持在不同网络中直接应用强化学习算法,无需大量重训练。为求解该模型,提出一种基于近端策略优化(PPO)的异步多智能体强化学习算法(AMAPPO)。
原文摘要 · Abstract (English)
The thesis proposes a generalized charging framework for multiple mobile chargers to maximize the network lifetime and ensure target coverage and connectivity in large scale WRSNs. Moreover, a multi-point charging model is leveraged to enhance charging efficiency, where the MC can charge multiple sensors simultaneously at each charging location. The thesis proposes an effective Decentralized Partially Observable Semi-Markov Decision Process (Dec POSMDP) model that promotes Mobile Chargers (MCs) cooperation and detects optimal charging locations based on realtime network information. Furthermore, the proposal allows reinforcement algorithms to be applied to different networks without requiring extensive retraining. To solve the Dec POSMDP model, the thesis proposes an Asynchronous Multi Agent Reinforcement Learning algorithm (AMAPPO) based on the Proximal Policy Optimization algorithm (PPO).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。