用深度强化学习让无人机群在复杂环境里自主追击目标,无需外部信息。
Decentralized End-to-End Multi-AAV Pursuit Using Predictive Spatio-Temporal Observation via Deep Reinforcement Learning
- 用预测时空观测网格直接处理激光雷达数据,实现端到端控制。
- 在仿真中捕获效率更高,且不同规模机群无需重训。
- 真实室外实验验证了仅靠机载感知和计算即可完成协同追击。
在杂乱环境中实现去中心化的自主飞行群协同追击极具挑战性,尤其在感知部分且噪声干扰下。现有方法常依赖抽象几何特征或特权真实状态,因而回避了真实场景中的感知不确定性。本文提出一种去中心化端到端多智能体强化学习(MARL)框架,将原始激光雷达观测直接映射为连续控制指令。核心是预测时空观测(PSTO),一种以自身为中心的网格表示,统一对齐障碍物几何、预测敌方意图及队友运动。基于PSTO,单一去中心化策略使智能体能避开静态障碍、拦截动态目标并维持协同包围。仿真结果表明,该方法在捕获效率上优于依赖特权障碍信息的先进学习方法,且策略可无缝扩展至不同团队规模而无需重训。此外,完全自主的室外实验验证了该框架在仅依赖机载感知与计算的四旋翼集群上的可行性。
原文摘要 · Abstract (English)
Decentralized cooperative pursuit in cluttered environments is challenging for autonomous aerial swarms, especially under partial and noisy perception. Existing methods often rely on abstracted geometric features or privileged ground-truth states, and therefore sidestep perceptual uncertainty in real-world settings. We propose a decentralized end-to-end multi-agent reinforcement learning (MARL) framework that maps raw LiDAR observations directly to continuous control commands. Central to the framework is the Predictive Spatio-Temporal Observation (PSTO), an egocentric grid representation that aligns obstacle geometry with predictive adversarial intent and teammate motion in a unified, fixed-resolution projection. Built on PSTO, a single decentralized policy enables agents to navigate static obstacles, intercept dynamic targets, and maintain cooperative encirclement. Simulations demonstrate that the proposed method achieves superior capture efficiency and competitive success rates compared to state-of-the-art learning-based approaches relying on privileged obstacle information. Furthermore, the unified policy scales seamlessly across different team sizes without retraining. Finally, fully autonomous outdoor experiments validate the framework on a quadrotor swarm relying on only onboard sensing and computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。