用AI智能调度无人机物流与计算任务,提升效率与可靠性。
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing

- 结合大模型与思维链,将用户指令转化为可执行的数学规划
- 分层强化学习使无人机在500轮中99.6%完成货物收集,100%满足任务截止时间
- 适合工业物联网中需协同调度物流与计算资源的场景
在云制造中,无人机可同时支持产品收集与移动边缘计算(MEC)。二者联合形成混合调度问题,物理物流决策与计算任务调度相互耦合。无人机从制造站点收集成品并运回中心仓库,同时处理站点传感器生成的计算任务——可在本地、无人机上或通过无人机卸载至云端。由于无人机仅在服务窗口内提供MEC能力,路径选择直接影响任务卸载时机。路径还影响无人机能耗、机载计算与通信资源可用性,需在任务截止时间内完成。为此,本文提出一种基于代理式AI的优化框架:第一,设计融合大语言模型、检索增强生成与思维链推理的代理系统,将用户输入转化为可解释的数学建模;第二,采用基于近端策略优化(PPO)的分层深度强化学习,上层学习无人机路径,下层优化每时隙任务执行与资源分配。仿真结果表明,该框架生成形式更一致,分层PPO在最后500轮中实现99.6%的产品完整收集,保持100%截止时间满足率,性能稳定性优于优势演员-评论家方法。
原文摘要 · Abstract (English)
In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with computational task scheduling. In this paper, UAVs collect finished products from manufacturing stations and transport them back to a central depot. Meanwhile, computational tasks generated by industrial sensor devices at these stations are processed locally, at UAVs, or offloaded via UAVs to the cloud. This coupling makes the problem challenging. A UAV can provide MEC services only during its service window at a station, so routing decisions directly determine when UAV-assisted offloading is available. Routing decisions also affect the UAV energy budget and the availability of onboard computing and communication resources for computational task execution under task deadline constraints. To address this, we propose an agentic-AI-assisted optimization framework with two components. First, we develop an agentic AI that combines large language models, retrieval-augmented generation, and chain-of-thought reasoning to translate user input into an interpretable mathematical formulation for the hybrid scheduling problem. Second, we design a hierarchical deep reinforcement learning approach based on proximal policy optimization (PPO), where the upper layer learns UAV routing and the lower layer optimizes per-slot task execution and resource allocation. Simulation results show that the proposed framework yields more consistent formulations, while the hierarchical PPO achieves full product collection in 99.6% of the last 500 episodes and maintains a 100% deadline satisfaction rate, with more stable performance than the advantage actor-critic approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。