arXiv:2605.26471cs.RO2026-05

用强化学习优化无人机物流任务分组,提升动态调度效率。

Heterogeneous AAV Logistics Task Allocation: A Reinforcement Learning Enhanced Overlapping Coalition Formation Game Approach

  • 结合强化学习与联盟形成博弈,动态生成高效重叠任务组
  • 32架无人机80个任务下,成本比传统方法降低39.76%
  • 适合城市空中物流、动态任务分配场景的智能系统研发者

在动态城市物流中,时效性任务随机出现给异构先进航空器(AAV)的任务分配带来显著优化挑战。为此,提出一种强化学习增强的重叠联盟形成博弈方法。构建动态任务分配模型,以广义物流成本(耦合服务质量与资源消耗)量化全局最优性。针对随机订单到达引发的时变任务集,设计基于Transformer的软演员-评论家网络,利用多头自注意力编码可变长度物流状态,捕捉任务级时空依赖关系,使学习策略自适应引导联盟更新,替代启发式规则。所提联盟形成过程被证明为精确势博弈,保证在有限步内收敛至纳什稳定均衡。数值仿真显示,该算法在32架AAV与80个任务场景下,相较启发式重叠联盟形成基线,实现39.76%的成本降低。室内飞行实验进一步验证其实际可行性。

原文摘要 · Abstract (English)

In dynamic urban logistics, the stochastic emergence of time-sensitive tasks poses a significant optimality challenge for heterogeneous AAVs logistics task allocation. To address this problem, a reinforcement learning enhanced overlapping coalition formation game approach is proposed. A dynamic task allocation model is established, where global optimality is mathematically quantified by a generalized logistics cost coupling service quality and resource consumption. To deal with the time-varying task sets induced by stochastic order arrivals, a transformer-based soft actor-critic network is designed. By leveraging multi-head self-attention to encode variable-length logistics states and capture task-wise spatiotemporal dependencies, the learned policy adaptively guides coalition updates, replacing heuristic rules in the overlapping coalition formation game. On this basis, heterogeneous AAVs can form more efficient overlapping coalitions for dynamic logistics tasks. The resulting coalition formation process is proven to constitute an exact potential game, which guarantees convergence to a Nash-stable equilibrium within a finite number of iterations. Numerical simulations demonstrate that the proposed algorithm effectively improves the optimality of task allocation under the generalized logistics cost criterion. In a scenario with 32 AAVs and 80 tasks, our algorithm achieves a 39.76% cost reduction compared with the heuristic OCF baseline. Indoor flight experiments further validate its practicality.

无人机物流强化学习任务分配博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。