用强化学习让不同公司无人机群自动避撞,还能自适应弱配置机型。
Separation Assurance between Heterogeneous Fleets of Small Unmanned Aerial Systems via Multi-Agent Reinforcement Learning

- 用增强注意力的PPOA2C算法,各无人机群独立训练策略保隐私
- 两种策略可达成安全平衡,强配置群体在博弈中占优
- 适合研究城市无人机调度与公平性问题的研究者
未来密集的城市空域中,多家公司运营包含多个同构无人机的异构机群,使战术避撞变得极为复杂。本文研究多智能体强化学习框架下,异构机群能否实现无冲突的运行均衡,以及强弱配置机群是否会被不公平对待。在德克萨斯州达拉斯进行模拟,采用基于注意力机制的PPOA2C框架,使各机群独立训练策略并保护隐私。实验表明,两组具有差异性的共享PPOA2C策略可达成安全均衡;相比两个强规则基基准,PPOA2C在冲突解决上表现更优,且与规则基策略交互更安全,体现其自适应能力。进一步评估显示,相似策略类型间均衡偏向配置更强的机群;即使配置相似,异构策略仍会形成偏向,凸显异构无人机运行中需引入公平性考量。
原文摘要 · Abstract (English)
In the envisioned future dense urban airspace, multiple companies will operate heterogeneous fleets of small unmanned aerial systems (sUASs), where each fleet includes several homogeneous aircraft with identical policies and configurations, e.g., equipage, sensing, and communication ranges, making tactical deconfliction highly complex for the aircraft. This paper aims to address two core questions: (1) Can tactical deconfliction policies converge or reach an equilibrium to ensure a conflict-free airspace when companies operate heterogeneous fleets of homogeneous aircraft? (2) If so, will the converged policies discriminate against companies operating sUASs with weaker configurations? We investigate a multi-agent reinforcement learning paradigm in which homogeneous aircraft within heterogeneous fleets operate concurrently to perform package delivery missions over Dallas, Texas, USA. An attention-enhanced Proximal Policy Optimization-based Advantage Actor-Critic (PPOA2C) framework is employed to resolve intra- and inter-fleet conflicts, with each fleet independently training its own policy while preserving privacy. Experimental results show that two fleets with distinct, shared PPOA2C policies can reach an equilibrium to maintain safe separation. While two PPOA2C policies outperform two strong rule-based baselines in terms of conflict resolution, a PPOA2C policy exhibits safer interaction with a rule-based policy, indicating adaptive capabilities of PPOA2C policies. Furthermore, we conducted extensive policy-configuration evaluations, which reveal that equilibria between similar policy types tend to favor fleets with stronger configurations. Even under similar configurations but different policy types, the equilibrium favors one of the heterogeneous policies, underscoring the need for fairness-aware conflict management in heterogeneous sUAS operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。