用前瞻风险引导强化学习,让无人机在动态障碍中更安全高效飞行
Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter

- 基于最近接近点构建方向对齐的未来碰撞风险图
- 在模拟与真实场景中均显著提升飞行安全与效率
- 无需物体追踪,直接从深度序列学习运动线索,适合物理飞行器部署
在复杂动态环境中实现四旋翼安全飞行,不仅依赖瞬时几何感知,更关键在于预判相对运动引发的碰撞风险。传统模块化流程常受感知延迟影响,而依赖隐式标量奖励的端到端方法往往缺乏物理约束,难以提取可靠的时空特征。为此,本文提出一种前瞻风险引导的强化学习框架。利用特权仿真状态,基于最近接近点(CPA)构建方向对齐的未来碰撞风险图。通过非对称的演员-评论家架构,网络自预测该结构化风险,在部署时显式指导视觉策略。轻量级时空编码器直接从机载深度序列提取运动线索,无需显式物体跟踪或光流估计。大量模拟与真实世界实验表明,相比现有基线,本方法在密集动态障碍中显著提升安全裕度与飞行效率。此外,所学策略在物理四旋翼上实现了零样本的仿真到现实迁移,仅依赖抽象化的时空深度序列及其自预测风险先验,验证了方法的有效性与强泛化能力。
原文摘要 · Abstract (English)
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。