arXiv:2602.01131cs.AI2026-02被引 1

用博弈论与稳定性理论,让无人机在资源有限时仍能稳定运行。

Lyapunov Stability-Aware Stackelberg Game for Low-Altitude Economy: A Control-Oriented Pruning-Based DRL Approach

  • 构建感知-通信-计算-控制闭环,用李雅普诺夫理论量化稳定性约束
  • 设计分层博弈机制,无人机动态定价资源,用户按紧急程度响应
  • 提出轻量剪枝PPO算法,训练时压缩模型,适配边缘设备实时决策

随着低空经济快速发展,无人机作为关键空中基站,支撑着从低延迟任务到高带宽流媒体等多种服务。然而,机载资源有限与严格稳定性要求之间的矛盾常导致系统效能下降。本文提出一个感知-通信-计算-控制闭环框架,显式建模通信延迟对物理控制稳定性的影响。基于李雅普诺夫稳定性理论,推导出控制系统状态演化与通信约束之间的内在映射,将抽象的稳定性需求转化为可量化的资源边界。随后将资源分配问题建模为斯塔克尔伯格博弈:无人机(领导者)动态定价以平衡负载并保障稳定性,用户(追随者)根据服务紧迫性优化请求。针对标准深度强化学习在能量受限边缘平台上的计算开销过大问题,提出一种新型轻量级剪枝式近端策略优化(PPO)算法。通过集成动态结构化剪枝机制,该算法在训练过程中显著压缩神经网络规模,使无人机能以极低推理延迟快速逼近博弈均衡。仿真结果表明,所提方案在动态低空环境中既能有效保障控制回路稳定性,又能最大化系统效用。

原文摘要 · Abstract (English)

With the rapid expansion of the low-altitude economy, Unmanned Aerial Vehicles (UAVs) serve as pivotal aerial base stations supporting diverse services from users, ranging from latency-sensitive critical missions to bandwidth-intensive data streaming. However, the efficacy of such heterogeneous networks is often compromised by the conflict between limited onboard resources and stringent stability requirements. Moving beyond traditional throughput-centric designs, we propose a Sensing-Communication-Computing-Control closed-loop framework that explicitly models the impact of communication latency on physical control stability. To guarantee mission reliability, we leverage the Lyapunov stability theory to derive an intrinsic mapping between the state evolution of the control system and communication constraints, transforming abstract stability requirements into quantifiable resource boundaries. Then, we formulate the resource allocation problem as a Stackelberg game, where UAVs (as leaders) dynamically price resources to balance load and ensure stability, while users (as followers) optimize requests based on service urgency. Furthermore, addressing the prohibitive computational overhead of standard Deep Reinforcement Learning (DRL) on energy-constrained edge platforms, we propose a novel and lightweight pruning-based Proximal Policy Optimization (PPO) algorithm. By integrating a dynamic structured pruning mechanism, the proposed algorithm significantly compresses the neural network scale during training, enabling the UAV to rapidly approximate the game equilibrium with minimal inference latency. Simulation results demonstrate that the proposed scheme effectively secures control loop stability while maximizing system utility in dynamic low-altitude environments.

无人机强化学习博弈论稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。