通过可微分世界模型,高效优化无人机轨迹以减少任务时长。
Learn for Variation: Efficient AAV Trajectory Learning through a Differentiable Wireless World Model
- 构建可微分世界模型,联合建模无人机运动、信道状态与用户队列。
- 采用累积队列替代完成时间目标,实现路径敏感度精确传播。
- 在多种复杂场景下表现优越,适合需快速部署的智能飞行系统。
自主空中载具(AAV)为第六代物联网网络提供数据采集支持,但其轨迹需协调非线性无线速率与长时程服务进度。本文将无人机运动学、信道状态和用户队列演化建模为结构化可微分世界模型,并提出学习变差(L4V)方法以高效利用该模型。L4V用累积队列代理替代不连续的完成时间目标,对任务动态进行展开,并通过离散伴随递推将路径敏感度反向传播至神经策略。所得导数在固定外生噪声实现实例下为精确值;随机期望目标优化仍需采样。我们证明结构伴随的增长最多为多项式级,且在标准光滑性假设下,固定步长全梯度下降具有平稳点收敛率。该框架还学习共享的OFDMA资源分配,在重参数化阴影与Rician衰落条件下运行,分布预训练使基于模型的优化转化为仅前向部署,适用于未知布局。压力测试涵盖信道生成器失配、部分观测噪声、固定资源双无人机扩展及圆形禁飞区。代码与配置见https://github.com/UNIC-Lab/L4V-AAV。相比遗传算法、DQN、A2C、DDPG及可微模型预测控制,L4V将任务时长减少高达65%,默认任务执行耗时仅53毫秒,且在预训练1,600个布局后完成全部60个冻结策略测试。
原文摘要 · Abstract (English)
Autonomous aerial vehicles (AAVs) enable data collection for sixth-generation Internet-of-Things networks, but their trajectories couple nonlinear wireless rates with long-horizon service progress. This paper views the evolution of AAV kinematics, channel state, and user backlog as a structured differentiable world model and develops Learn for Variation (L4V) to exploit that model efficiently. L4V replaces a discontinuous completion-time objective with a cumulative-backlog surrogate, unrolls the mission dynamics, and propagates pathwise sensitivities to a neural policy through the discrete adjoint recursion. The resulting derivative is exact conditional on a fixed exogenous-noise realization; stochastic expected-objective optimization still requires sampling. We show that the structured adjoint grows at most polynomially with the horizon and establish a stationary-point rate for fixed-step full-gradient descent under standard smoothness assumptions. The framework also learns shared OFDMA allocation under reparameterized shadowing and Rician fading, while distributional pretraining amortizes model-based optimization into forward-only deployment on unseen layouts. Paired stress tests cover channel-generator mismatch, noisy partial observations, a fixed-resource two-AAV extension, and a circular no-fly region. Code and configurations are available at https://github.com/UNIC-Lab/L4V-AAV. Against genetic-algorithm, DQN, A2C, DDPG, and differentiable model-predictive-control implementations, L4V reduces mission time by up to $65\%$, executes a default mission in $53$ ms, and completes all $60$ frozen-policy tests after pretraining on $1{,}600$ layouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。