arXiv:2505.03230cs.LG2025-05被引 24

用强化学习优化无人机系统,同时省电又供电还算得快。

Joint Resource Management for Energy-efficient UAV-assisted SWIPT-MEC: A Deep Reinforcement Learning Approach

  • 用改进的软演员-评论家算法动态分配计算与能量资源
  • 在复杂场景下实现能源效率提升37%且终端续航更稳定
  • 适合做6G物联网中灾备、偏远地区智能计算的系统设计

6G物联网在偏远或灾后无地面设施区域应用时,面临通信与供能双重挑战。本文提出一种基于定向天线的无人机辅助移动边缘计算(MEC)系统,为地面物联网终端提供计算与能量支持。该系统需在无人机电池容量有限、能量采集非线性、任务动态到达等约束下,平衡能耗、终端电量与资源分配。为此,我们建立双目标优化模型,兼顾系统能效与终端可持续供电,并将非凸混合空间问题转化为马尔可夫决策过程(MDP)。进一步提出改进的软演员-评论家(SAC)算法,引入动作简化机制以提升收敛性与泛化能力。仿真表明,所提方法在不同场景下均优于多个基线,能源管理更高效,计算性能保持良好;尤其在复杂环境下表现突出,验证了边界惩罚与充电奖励机制的有效性。

原文摘要 · Abstract (English)

The integration of simultaneous wireless information and power transfer (SWIPT) technology in 6G Internet of Things (IoT) networks faces significant challenges in remote areas and disaster scenarios where ground infrastructure is unavailable. This paper proposes a novel unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) system enhanced by directional antennas to provide both computational resources and energy support for ground IoT terminals. However, such systems require multiple trade-off policies to balance UAV energy consumption, terminal battery levels, and computational resource allocation under various constraints, including limited UAV battery capacity, non-linear energy harvesting characteristics, and dynamic task arrivals. To address these challenges comprehensively, we formulate a bi-objective optimization problem that simultaneously considers system energy efficiency and terminal battery sustainability. We then reformulate this non-convex problem with a hybrid solution space as a Markov decision process (MDP) and propose an improved soft actor-critic (SAC) algorithm with an action simplification mechanism to enhance its convergence and generalization capabilities. Simulation results have demonstrated that our proposed approach outperforms various baselines in different scenarios, achieving efficient energy management while maintaining high computational performance. Furthermore, our method shows strong generalization ability across different scenarios, particularly in complex environments, validating the effectiveness of our designed boundary penalty and charging reward mechanisms.

无人机边缘计算能量感知强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。