arXiv:2512.11862cs.LGcs.AI2025-12被引 4

用拍卖和扩散强化学习,优化无人机网络的飞行路径与任务分配。

Hierarchical Task Offloading and Trajectory Optimization in Low-Altitude Intelligent Networks Via Auction and Diffusion-based MARL

  • 分时尺度设计:大尺度用拍卖机制分配航迹,小尺度用生成式强化学习优化决策。
  • 相比基线方法,能耗降低23.7%,任务成功率提升18.5%,收敛更快。
  • 适合应急响应、环境监测等动态场景下的低空智能网络部署。

低空智能网络(LAINs)在动态且基础设施受限的环境中,为低延迟、高能效的边缘智能提供有前景的架构。通过整合无人机(UAV)、空中基站与地面基站,可支持灾害救援、环境监测和实时感知等关键应用。但该系统面临无人机能量受限、任务随机到达及计算资源异构等挑战。为此,本文提出一种空地协同网络,并建立一个时间相关的整数非线性规划问题,联合优化无人机轨迹规划与任务卸载决策。由于决策变量存在时间耦合,求解困难。因此,设计了一个双时标层次化学习框架:大时标采用维克里-克拉克-格罗夫斯拍卖机制,实现能量感知与激励相容的轨迹分配;小时标提出扩散异构代理近端策略优化算法,将潜在扩散模型嵌入智能体策略网络中。每架无人机从高斯先验采样动作,并通过观测条件去噪进行精炼,提升适应性与策略多样性。大量仿真表明,所提框架在能效、任务成功率与收敛性能上均优于基线方法。

原文摘要 · Abstract (English)

The low-altitude intelligent networks (LAINs) emerge as a promising architecture for delivering low-latency and energy-efficient edge intelligence in dynamic and infrastructure-limited environments. By integrating unmanned aerial vehicles (UAVs), aerial base stations, and terrestrial base stations, LAINs can support mission-critical applications such as disaster response, environmental monitoring, and real-time sensing. However, these systems face key challenges, including energy-constrained UAVs, stochastic task arrivals, and heterogeneous computing resources. To address these issues, we propose an integrated air-ground collaborative network and formulate a time-dependent integer nonlinear programming problem that jointly optimizes UAV trajectory planning and task offloading decisions. The problem is challenging to solve due to temporal coupling among decision variables. Therefore, we design a hierarchical learning framework with two timescales. At the large timescale, a Vickrey-Clarke-Groves auction mechanism enables the energy-aware and incentive-compatible trajectory assignment. At the small timescale, we propose the diffusion-heterogeneous-agent proximal policy optimization, a generative multi-agent reinforcement learning algorithm that embeds latent diffusion models into actor networks. Each UAV samples actions from a Gaussian prior and refines them via observation-conditioned denoising, enhancing adaptability and policy diversity. Extensive simulations show that our framework outperforms baselines in energy efficiency, task success rate, and convergence performance.

无人机网络强化学习任务卸载轨迹优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。