arXiv:2603.10528cs.LGcs.AI2026-03被引 3

用强化学习协调无人机群,实时高效配送紧急医疗物资。

UAV-MARL: Multi-Agent Reinforcement Learning for Time-Critical and Dynamic Medical Supply Delivery

  • 基于PPO的多智能体强化学习框架,动态调度无人机资源。
  • 在真实地理数据上测试,经典PPO比异步策略表现更优。
  • 适合应急医疗物流、无人机调度系统研究者参考。

无人机(UAV)在紧急情况和资源短缺时日益用于快速灵活地配送关键医疗物资。然而,有效部署无人机机群需要能够优先处理医疗请求、分配有限空中资源,并在不确定条件下动态调整配送计划的协调机制。本文提出一种多智能体强化学习(MARL)框架,用于应对随机性医疗配送场景中请求紧急程度、位置和交付截止时间各异的问题。该问题被建模为部分可观测马尔可夫决策过程(POMDP),其中无人机智能体需在通信与定位受限导致视野有限的情况下感知医疗需求。所提框架采用近端策略优化(PPO)作为主算法,评估了异步扩展、经典演员-评论家方法及架构改进等多种变体,以分析可扩展性与性能权衡。模型基于从OpenStreetMap提取的真实医疗机构地理数据进行评估。结果表明,经典PPO在协调性能上优于异步与串行学习策略,凸显强化学习在自适应、可扩展的无人机辅助医疗物流中的潜力。

原文摘要 · Abstract (English)

Unmanned aerial vehicles (UAVs) are increasingly used to support time-critical medical supply delivery, providing rapid and flexible logistics during emergencies and resource shortages. However, effective deployment of UAV fleets requires coordination mechanisms capable of prioritizing medical requests, allocating limited aerial resources, and adapting delivery schedules under uncertain operational conditions. This paper presents a multi-agent reinforcement learning (MARL) framework for coordinating UAV fleets in stochastic medical delivery scenarios where requests vary in urgency, location, and delivery deadlines. The problem is formulated as a partially observable Markov decision process (POMDP) in which UAV agents maintain awareness of medical delivery demands while having limited visibility of other agents due to communication and localization constraints. The proposed framework employs Proximal Policy Optimization (PPO) as the primary learning algorithm and evaluates several variants, including asynchronous extensions, classical actor--critic methods, and architectural modifications to analyze scalability and performance trade-offs. The model is evaluated using real-world geographic data from selected clinics and hospitals extracted from the OpenStreetMap dataset. The framework provides a decision-support layer that prioritizes medical tasks, reallocates UAV resources in real time, and assists healthcare personnel in managing urgent logistics. Experimental results show that classical PPO achieves superior coordination performance compared to asynchronous and sequential learning strategies, highlighting the potential of reinforcement learning for adaptive and scalable UAV-assisted healthcare logistics.

无人机配送强化学习医疗物流多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。