用可微分旋翼机动力学学习高速追击策略,无需精确位置信息。
Learning Agile Intruder Interception using Differentiable Quadrotor Dynamics

- 通过3D方向向量和拦截器状态设计控制策略
- 在10米/秒速度下性能比基线方法提升30%
- 适合单目摄像头等受限感知场景的无人机追击任务
本文提出一种基于可微分旋翼机动力学的学习方法,利用到入侵者的目标3D方向单位向量与拦截器状态来学习控制策略。以往深度强化学习方法依赖相对位置或距离信息,但这些信息在使用被动单目摄像头的实际应用中难以获取。本文提出的方法采用解析式策略梯度法,在不依赖精确位置的前提下,实现了最高达10米/秒的敏捷拦截。相比使用简化质点动力学的基线方法,该方法平均性能提升30%。
原文摘要 · Abstract (English)
This paper presents a methodology for learning a control policy to intercept an intruder using the 3D direction unit vector to the intruder and the interceptor state. Prior deep reinforcement learning approaches assume either relative position or distance to the intruder is available, but this information is not readily accessible in real-world applications that employ passive, monocular camera sensors. Instead, we propose a solution that leverages an analytical policy gradient method using differentiable quadrotor dynamics to learn agile interception at speeds up to 10 m/s. The proposed approach outperforms baseline methods that utilize simplified point mass dynamics by an average of 30%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。