arXiv:2608.14135cs.ROcs.LG2026-08

无人机自主追逃博弈,用自对弈强化学习实现零样本真实飞行

AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning

论文配图:AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning
图 1 · 摘自论文原文
  • 自对弈强化学习直接输出飞行控制指令,无需中间路径规划
  • 在仿真中训练出复杂追逃策略,真实四轴飞行器零样本部署成功
  • 模拟环境精准复现硬件延迟与扰动,适合无人系统实时决策研究

自主追逃是无人机面临的根本挑战,要求在高度耦合的动力学和不断变化的对手行为下快速决策。传统规则或微分博弈方法难以应对高维空域交互和敏捷机动。我们提出AgilePE,一种基于自对弈强化学习的完整无人机自主追逃系统。该系统集成敏捷低层控制、竞争性策略优化与仿真到现实的部署框架。策略直接将机载状态观测映射为集体推力与机体角速率(CTBR)指令,实现端到端敏捷机动,无需中间轨迹规划或航点控制器。训练采用优先虚构自对弈(PFSP)和多样化对手池,使智能体在对抗历史策略的同时稳定优化,减少策略振荡,从而涌现出复杂追逃策略。针对真实部署,我们构建了硬件对齐的仿真流水线,建模执行器响应动态、通信延迟和领域随机化。所学策略可零样本迁移至真实四旋翼无人机,无需任务特化调优。真实实验再现了仿真中的追逃战术,包括快速闪避与包抄,并验证了双智能体零样本交互部署能力。

原文摘要 · Abstract (English)

Autonomous pursuit-evasion is a fundamental challenge for Unmanned Aerial Vehicles (UAVs), requiring rapid decision-making under tightly coupled dynamics and continuously changing opponent behaviors. Traditional rule-based or differential-game approaches often struggle with high-dimensional aerial interactions and agile maneuvering. We present AgilePE, a complete system for autonomous UAV pursuit-evasion via self-play reinforcement learning. AgilePE integrates agile low-level control, competitive policy optimization, and sim-to-real deployment in a unified framework. The policy directly maps onboard state observations to Collective Thrust and Body Rates (CTBR) commands, enabling end-to-end agile maneuvering without intermediate trajectory planners or waypoint controllers. For training, we use competitive self-play with Prioritized Fictitious Self-Play (PFSP) and a diversified opponent pool, enabling agents to improve against historical policies while stabilizing optimization and reducing policy oscillation. This process leads to the emergence of sophisticated pursuit and evasion strategies. For real-world deployment, we develop a hardware-aligned simulation pipeline that models actuator-response dynamics, communication latency, and domain randomization. The learned policies transfer zero-shot to real quadrotors without task-specific tuning. Real-world experiments reproduce pursuit-evasion tactics observed in simulation, including rapid dodging and flanking, and demonstrate interactive two-agent zero-shot deployment.

无人机强化学习自对弈零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。